This paper aims to investigate the relationship between expert and consumer wine ratings and their association with wine prices. Using a unique longitudinal data set of an informal wine club, it explores whether these two forms of evaluation align or diverge and how each group perceived wine quality via price premiums.
The analysis, spanning 1987–2024, is based on 693 paired ratings from two distinct sources: expert evaluations as published in Platter’s South African Wine Guide, and consumer evaluations from an informal but experienced private South African wine club. Agreement between these two rating sources is tested using Cohen’s weighted kappa statistic. A Bayesian regression model is used to assess how each rating source relates to real (consumer price index-deflated) retail wine prices.
Although the expert and wine club scores show very low agreement, both independently predict wine prices. Expert ratings have a stronger effect, but consumer ratings also show significant predictive power. This relationship hold across both red and white wines.
Wine producers and marketers can benefit from displaying both expert and consumer scores, as each appeals to different segments of the market. The data set also underscores the value of long-term consumer panels in understanding price sensitivity.
This study presents a novel data set that enables the examination of nearly four decades of matched expert and enthusiast wine evaluations in South Africa. It demonstrates how different evaluative perspectives coexist in shaping price premiums, even in the face of diverging views on quality.
1. Introduction
Wine selection is a complex decision-making process influenced by a multitude of factors, including sensory preferences, expert evaluations and price. As global wine markets have grown increasingly diverse, consumers have come to rely more on external cues, such as expert ratings, to guide their choices. These ratings, frequently awarded during tasting competitions and specialist publications, are widely regarded as indicators of quality. However, the extent to which they align with general consumer preferences remains debated.
Previous research, such as the study by Schiefer and Fischer (2008), highlights the gap between wine expert ratings and consumer preferences, emphasizing that expert evaluations may not reliably predict the sensory enjoyment of the average consumer. Moreover, consumer preferences tend to be heterogeneous, further complicating the development of universally applicable quality indicators. Schiefer and Fischer (2008)’s findings suggest that while experienced and knowledgeable wine drinkers might align more closely with expert opinions, the broader consumer base may benefit more from alternative approaches, such as sensory descriptors tailored to their tastes.
Building on this foundation, this study examines the relationship between aggregated consumer ratings and expert ratings in a unique and novel panel data set spanning close to four decades. In addition to these ratings, the analysis incorporates factors such as wine price, type (red or white) and the year of the tasting to explore the determinants of wine preference and evaluate the alignment between consumer perceptions and expert evaluations. This contextual approach provides a comprehensive understanding of the dynamics influencing wine preference and quality perception, contributing to ongoing discussions about the practical value of expert ratings in guiding consumer decisions.
2. Literature review
The relationship between wine ratings and price premiums has long been a focal point in wine economics. In particular, the understanding of how expert evaluations influence market dynamics is a key question. In this context, Ali et al. (2008) provide compelling evidence of the strong influence of Robert Parker’s ratings on Bordeaux en primeur prices. Their analysis shows that even small increases in expert scores can lead to substantial price premiums (see also Babin and Bushardt, 2019). These findings underscore the powerful role of expert evaluations in shaping consumer perceptions and willingness-to-pay, often inflating prices beyond what intrinsic quality alone would warrant. That said, the degree to which other rating systems (such as consumer-driven reviews or alternative guides) affect price formation remains comparatively underexplored.
Expanding the analysis to the dynamic interaction between ratings and prices, Kaimann et al. (2023) investigate the interplay of quality signals in the wine market using a data set of 13,911 observations over two decades. They demonstrate that expert ratings significantly influence prices, with a one-point increase in rating associated with an 8% price rise. Lagged ratings also exert considerable influence on current scores, indicating path dependence in expert assessments. The study further reveals that expert ratings disproportionately impact high-quality wines, emphasizing the global relevance of these findings across diverse regions and varieties. This aligns with earlier research highlighting the amplifying effects of positive reviews on price premiums for top-tier wines (Villas-Boas et al., 2021; Oczkowski and Pawsey, 2019; Ashton, 2016).
Goldstein et al. (2008) present a pivotal contribution by analysing over 6,000 blind tastings conducted across a diverse sample of participants in the USA. Their study challenges the commonly held assumption that more expensive wines inherently deliver greater enjoyment, revealing a nuanced relationship contingent upon the taster’s level of expertise. For non-expert consumers, there exists a small but statistically significant negative correlation between price and subjective enjoyment, while expert tasters exhibit a non-negative relationship, suggesting that expertise aligns perceptions more closely with premium wine attributes. This finding underscores the limitations of price as a signal of quality for non-experts and highlights the psychological influence of extrinsic cues such as pricing and branding on perceived quality (Plassmann et al., 2008).
Extending this inquiry, Chocarro and Cortinas (2013) examine how expert opinions shape consumer evaluations of wine, emphasizing the moderating effects of information complexity, expert consensus and consumer knowledge. Their experimental design demonstrates that expert ratings significantly influence evaluations, particularly among consumers with low wine knowledge. This group relies on these “weak-tie” sources to reduce uncertainty when evaluating experience goods. By contrast, high-knowledge consumers exhibit greater independence in their evaluations, likely due to their confidence in their own judgement. Interestingly, neither expert consensus nor information complexity significantly affected consumer responses, suggesting a preference for straightforward and accessible ratings. The study highlights the segmentation within the wine market, where differing levels of consumer knowledge necessitate tailored marketing strategies.
Building on this foundation, Oczkowski and Doucouliagos (2015) provide a comprehensive meta-regression analysis of 180 hedonic wine price models spanning two decades and multiple countries. They identify a moderate partial correlation of +0.30 between price and quality ratings, emphasizing the dual role of intrinsic attributes, such as sensory quality, and extrinsic factors, such as producer reputation, in shaping wine prices. Reputation often dominates as a price determinant, yet sensory quality retains a statistically significant impact, particularly when fine-grained quality measures like 100-point scales are used. Their findings demonstrate the value of methodological rigour in uncovering systematic heterogeneity across studies, further broadening the understanding of price–quality relationships in diverse contexts. Where Goldstein et al. (2008) reveal a disconnect between price and enjoyment for non-experts, Oczkowski and Doucouliagos (2015) and Oczkowski (2018) emphasize the broader interplay of sensory quality and reputation, extending the scope of inquiry by analysing how methodological design influences results.
Their analysis is supported by Núñez et al. (2024), who perform a meta-regression of 223 articles estimating hedonic functions and document selection and publication biases. They find robust price elasticities around 5%–6%; quality ratings have a larger impact on price elasticity for red wines; the type of data source matters, but the econometric techniques used do not. The remaining attributes mostly have small influences. Interestingly, the most cited studies are those that show a smaller effect.
Turning to the cognitive dimensions of wine evaluation, Masset and Raub (2023) examine the influence of extrinsic cues − such as price, expert ratings and bottle presentation − on wine quality perceptions and willingness-to-pay. Their experimental study, building on the seminal work of Hodgson (2008), highlights how tasters often fail to identify identical wines served under different conditions, favouring those with higher-status cues. Interestingly, experienced tasters adapt their ratings more strongly to extrinsic information, driven by confidence in their expertise and a desire to align with normative expectations. This work underscores how cognitive biases influence wine evaluation, reiterating findings from earlier studies showing that price cues affect neural processing of taste (Plassmann et al., 2008) and that non-experts derive less enjoyment from expensive wines when blind to price information.
2.1 Wine consumer behaviour and purchase cues in South Africa
There is also ample South African evidence that extrinsic information shapes perceived quality and purchase cues over time.
Conceptually, Priilaid (2006) collates blind and sighted tastings of South African reds (1993–2001) and show that sighted quality ratings are dominated by extrinsic cues (incl. price and region of origin). Their meta-model showed an adjusted of , where price alone explains roughly 84% of variance in sighted assessments. As for the combination of price and origin, this explained about 95% of the variance. The results also showed that intrinsic merit contributes only marginally once the extrinsic cues were controlled for. This divergence between sighted and blind scores implies that extrinsic information can mask underlying quality signals. But what about other cues like awards?
Herbst and Von Arnim (2009) find that although South African consumers recognize awards as a quality cue, they are secondary to variety, vintage, producer reputation, origin and price. Their study showed that more knowledgeable consumers are relatively sceptical and support independent oversight of competitions. They do however note in their study that decision-making is a complex set of interactions, and wine awards do nevertheless play a role in supporting consumer decisions in certain circumstances and for certain customer segments. Generational segmentation further nuances the cue effects.
In a cohort-focused study of South African Generation Y (Gen Y) consumers, Lategan et al. (2017) use best–worst scaling (BWS) on 13 attributes and report that “tasted previously” () and “someone recommended it” () are dominant cues. For Gen Y, personal experience and word of mouth in wine selection is most important. Extrinsic heuristics also feature: “brand” (), “in-store promotion” () and “medal/award” (). These results are consistent with the use of simplified quality signals among less knowledgeable consumers [1]. In a broader study that also used BWS across 14 attributes, and featured a national panel (), Pentz and Forrester (2020) likewise find “previous tasting” and “peer recommendations” to be most influential, while “in-store promotions” are least valued by older consumers and “low alcohol” by younger consumers. Interestingly “award” featured dead in the middle with a , but had a standard deviation of . This latter result signals some respondents value awards, while others actively down-weight them. It is likely that awards act as a secondary cue, and in this study, the individuals use it to support primary signals such as familiarity (prior tasting) and social proof (word of mouth).
The data presented in the next section enable us to examine the “gurus vs geeks” questions that we later empirically test: (1) contrasting expert and community signals (Bazen et al., 2024; Ali et al., 2008; Friberg and Grönqvist, 2012; Oczkowski and Pawsey, 2019), and to (2) quantify the rating–price linkage and its dynamics (Lecocq and Visser, 2006; Oczkowski and Doucouliagos, 2015; Kaimann et al., 2023; Núñez et al., 2024).
3. About the club
A data set was obtained from a wine club that has been in existence since 1987, with the name/producer, cultivar or blend, price, average score and ranking (within each tasting) recorded in a database for most of the wines ever tasted [2].
The wine club consists of 12 couples and strives to hold 12 blind tastings per year (i.e. one per month), although holidays and other factors usually result in around 10 tastings annually. Hosting duties rotate equally among the couples. Each couple has carte blanche regarding the topic, venue and presentation of their tasting. The majority of tastings are hosted at members’ homes, although some have taken place in hired venues, restaurants or the wine farms themselves.
The club’s membership has remained remarkably stable over its 37-year existence, with nine of the original couples still active. The composition of the members is diverse. One is a Cape Wine Master who produces under their own label and five others are winemakers (one of whom is retired). These six members are regarded as expert tasters within the group. Almost all of the other male members are avid wine lovers with varying degrees of tasting ability. Most of the ladies also enjoy wine and, like the non-winemaker male members, have differing levels of tasting expertise. The average age of the club members is 62, meaning the group was in their late twenties when the club was formed.
The club uses a 20-point scoring system, nominally with 3 points for colour, 7 for nose and 10 for taste. However, this breakdown is not “officially” applied, and only the aggregate score (out of 20) is recorded for each wine. Average attendance at each tasting is around nine male and six female members, resulting in 15 scores contributing to each wine’s average. The format of nearly all tastings includes a welcome glass (of any kind), followed by the formal tasting (typically 6–8 wines), and then dinner.
Given the club’s demographic composition, it is important to note that gendered consumption differences may influence how tasters use information and form preferences − even in a blind tasting. Evidence reports systematic differences in sensory preference and behaviour between females and males [3]. These differences tend to dissipate with experience (Bruwer et al., 2011). Our data set does not include gender-coded scores, so we cannot test these mechanisms directly. We now note this as a limitation and an avenue for future research.
4. Data analysis
This section draws on two data sources:
an informal wine tasting club; and
the Platter’s South African Wine Guide.
The wine club, based in Stellenbosch − the epicentre of South Africa’s wine industry and a hub for wine tourism — has held regular tastings since its inception in August 1987. Over the period from 1987 to April 2024, members hosted 352 monthly tastings, evaluating a total of 2,460 wines.
For each tasting, records were kept of the host, date, theme and detailed information on each wine, including its origin, cultivar, price and the group’s average score (on a 20-point scale). While the data set is largely complete, in some instances price or score data were missing. The number of tastings conducted per year is illustrated in Figure 1 [4].
The histogram illustrates the count of wines per year from 1990 to 2020, focusing on wine club ratings. The x-axis represents years, marked at intervals from 1990 to 2020, while the y-axis indicates the number of wines with tick marks. Each vertical bar corresponds to a specific year, showing the quantity of wines rated, with varying heights indicating their respective counts. The bars display a range of heights, with some years showing higher counts, suggesting fluctuations in wine ratings over time.Number of wines tasted per year
Note(s): Bars show the total number of wines evaluated at formal tastings per calendar year. The dataset covers 37 years (1987–April 2024) with a total of 693 wines paired with Platter’s Guide ratings. Annual participation ranged from a minimum of three wines in 1999 to a maximum of 36 in 2009, with a median of 18. The interquartile range (11.8–26) reflects moderate variation in annual tasting intensity across the period
The histogram illustrates the count of wines per year from 1990 to 2020, focusing on wine club ratings. The x-axis represents years, marked at intervals from 1990 to 2020, while the y-axis indicates the number of wines with tick marks. Each vertical bar corresponds to a specific year, showing the quantity of wines rated, with varying heights indicating their respective counts. The bars display a range of heights, with some years showing higher counts, suggesting fluctuations in wine ratings over time.Number of wines tasted per year
Note(s): Bars show the total number of wines evaluated at formal tastings per calendar year. The dataset covers 37 years (1987–April 2024) with a total of 693 wines paired with Platter’s Guide ratings. Annual participation ranged from a minimum of three wines in 1999 to a maximum of 36 in 2009, with a median of 18. The interquartile range (11.8–26) reflects moderate variation in annual tasting intensity across the period
Some tastings included categories such as beer, fortified wines, sparkling wines, rosé and non-South African wines. In several cases, price information was also missing. We excluded all six of these categories from the database. Furthermore, we removed all wines for which the Club rating was not recorded. This step was necessary to ensure alignment with our second data source, the Platter’s South African Wine Guide.
Platter’s Guide is a unique annual South African publication that records every wine that is commercially available in the country each year. Wines are tasted during the course of the year in question by a panel of expert wine tasters, and are then rated on a fivestar scale. The tasters are sighted, i.e. tastings are not blind. Each wine in our final database therefore had two ratings, namely that of the Club and the Platter rating. The data set spans close to four decades, offering a rich panel component that allows for the investigation of trends and patterns over space and in time. This meant that the full sample of wines included in the analysis amounted to 693.
The paired wine Club and Platter’s ratings form the basis for the examination of the relationship between expert evaluations, pricing and consumer preferences. The analysis begins with a breakdown of observations per year, further categorized by wine type, offering a view of the data set’s time series characteristics. The distribution of wine types and vintages is then explored to provide insights into the diversity and representativeness of the data set. Finally, we examine the ratings variable, including its range, distribution and trends over time, which are integral to understanding the dynamics of expert evaluations in the wine market.
The variables recorded and used in the analysis are described in Table 1.
Summary of variables
| Variable | Description | Data sample |
|---|---|---|
| Date | The exact date of the observation in DD/MM/YYYY format | 28 / 04/2005 |
| Yearbook | The yearbook used for the Platter rating | 2005 |
| Wine | The name of the winery being rated | Simonsig |
| Cultivar type | The grape variety or type of the wine | Chenin Blanc |
| Type | The category of the wine (e.g. red, white, sparkling) | White |
| Vintage | The year the wine was produced (vintage year) | 2004 |
| Price | The listed price of the wine in nominal Rand terms | R22.00 |
| Average score (club) | The average rating assigned by the Club out of 20 | 15.5 |
| Average score (Platter’s) | The average rating assigned by Platter’s wine guide out of 5 | 2.5 |
| CPI | The consumer price index (CPI) for the observation period | 57.0 |
| Deflated price (CPI) | The price of the wine adjusted for inflation using the CPI | R57.70 |
| Deflated price (wine) | The wine price adjusted for inflation using a wine-specific deflator | R60.30 |
| Variable | Description | Data sample |
|---|---|---|
| Date | The exact date of the observation in DD/MM/YYYY format | 28 / 04/2005 |
| Yearbook | The yearbook used for the Platter rating | 2005 |
| Wine | The name of the winery being rated | Simonsig |
| Cultivar type | The grape variety or type of the wine | Chenin Blanc |
| Type | The category of the wine (e.g. red, white, sparkling) | White |
| Vintage | The year the wine was produced (vintage year) | 2004 |
| Price | The listed price of the wine in nominal Rand terms | R22.00 |
| Average score (club) | The average rating assigned by the Club out of 20 | 15.5 |
| Average score (Platter’s) | The average rating assigned by Platter’s wine guide out of 5 | 2.5 |
| The consumer price index ( | 57.0 | |
| Deflated price ( | The price of the wine adjusted for inflation using the | R57.70 |
| Deflated price (wine) | The wine price adjusted for inflation using a wine-specific deflator | R60.30 |
This table summarizes the variables used in the analysis along with a representative sample entry. Prices are expressed in South African Rand (ZAR). Inflation adjustments were made using both general CPI and a wine-specific deflator. The Club rating is based on the discussed 20-point scale, while Platter’s ratings follow its five-star format
The choice of wines that were tasted is quite even with 363 red and 330 white wines included in the data. The data also show that the preferences of the Club change over the course of the year. White wines were generally preferred during the summer months (January to March, but in their case stretching into April) and red wines during most of autumn (fall) and winter (in this case May to August), with some caveats during spring, when white wines were preferred during September to November, but red during December (Figure 2) [5].
This bar graph illustrates the monthly proportions of wine consumed throughout the year, segmented by red and white varieties. The x-axis lists the months from January to December, while the y-axis shows the proportion of wine per month in percentages from zero to one hundred. Each monthly bar is divided into two sections: the section representing the proportion of white wine, and another section indicating the proportion of red wine. The data is arranged vertically with distinct heights for each segment, highlighting the variations in wine preferences for each month.Proportion of wine tasted in a given month
Note(s): Stacked bars show the proportion of red and white wines evaluated per month across 693 paired tastings (1987–2023). Seasonal preferences are evident: white wines dominate in the summer months (January–March), while red wines predominate through autumn and winter (May–August). Spring shows a mixed profile, with whites slightly favoured in September–November. December tastings are fewer in number (red , white ). A chi-squared test confirms a statistically significant association between wine type and month (, df = 11, )
This bar graph illustrates the monthly proportions of wine consumed throughout the year, segmented by red and white varieties. The x-axis lists the months from January to December, while the y-axis shows the proportion of wine per month in percentages from zero to one hundred. Each monthly bar is divided into two sections: the section representing the proportion of white wine, and another section indicating the proportion of red wine. The data is arranged vertically with distinct heights for each segment, highlighting the variations in wine preferences for each month.Proportion of wine tasted in a given month
Note(s): Stacked bars show the proportion of red and white wines evaluated per month across 693 paired tastings (1987–2023). Seasonal preferences are evident: white wines dominate in the summer months (January–March), while red wines predominate through autumn and winter (May–August). Spring shows a mixed profile, with whites slightly favoured in September–November. December tastings are fewer in number (red , white ). A chi-squared test confirms a statistically significant association between wine type and month (, df = 11, )
Another temporal component to wine consumer preferences that is of interest, is the vintage of the wines tasted. Most reds, having different ageing potential dependent on the time in barrel fermentation and time in bottle. Figure 3 shows the estimated mean cellaring (or yearbook − vintage) with confidence intervals for red wines over time [6]. The figure illustrates the estimated marginal mean cellaring time per year over a 37-year period. There appears to be a slight increase in the cellaring duration for red wines before tasting. Although this is not the focus in this paper, the trend could reflect the Club’s evolving preferences as they gain access to, and develop a taste for older vintages, or it might indicate a broader change within the wine market.
This line graph illustrates the average difference between year and vintage from 1990 to 2023. The x-axis represents the years, while the y-axis indicates the average difference values, ranging from zero to a maximum of five. A line traces the average differences over the years, showing notable fluctuations and peaks. The graph includes another line representing a trend, alongside a shaded area that indicates confidence intervals surrounding the average line.Estimated mean with credible intervals for vintages over time
Note(s): The solid blue line shows the annual mean difference between tasting year and vintage (“cellaring time”), with shaded bands indicating 95% confidence intervals. The fitted red line is from an OLS regression of cellaring time on year. Results suggest a slight upward trend in the age of red wines at tasting. OLS regression diagnostics: residual standard error = 1.746 on 329 df; adjusted ; , . Estimated marginal means indicate that cellaring time varied between and 5.7 years across the sample. This pattern may reflect both evolving club preferences and wider market availability of older vintages
This line graph illustrates the average difference between year and vintage from 1990 to 2023. The x-axis represents the years, while the y-axis indicates the average difference values, ranging from zero to a maximum of five. A line traces the average differences over the years, showing notable fluctuations and peaks. The graph includes another line representing a trend, alongside a shaded area that indicates confidence intervals surrounding the average line.Estimated mean with credible intervals for vintages over time
Note(s): The solid blue line shows the annual mean difference between tasting year and vintage (“cellaring time”), with shaded bands indicating 95% confidence intervals. The fitted red line is from an OLS regression of cellaring time on year. Results suggest a slight upward trend in the age of red wines at tasting. OLS regression diagnostics: residual standard error = 1.746 on 329 df; adjusted ; , . Estimated marginal means indicate that cellaring time varied between and 5.7 years across the sample. This pattern may reflect both evolving club preferences and wider market availability of older vintages
Besides the type of wine being tasted within this novel data set, a unique aspect that was captured within the notes are the ratings and prices paid for the various wines. Ratings themselves are interesting and much research has been conducted in this space. Two primary methods are commonly used in wine tasting evaluations due to their simplicity: (1) score average [7], and (2) the rank average [8] (Ashenfelter and Quandt, 1999). Figure 4 shows the distributional comparison between the five-point Platters’s Guide ratings and the 20-point Club rating.
The image features two bar graphs side by side. The left graph presents Platter ratings on the horizontal axis, ranging from two to five, with the vertical axis displaying relative percentages from zero to thirty percent. The bars indicate the frequency of each rating, peaking at four. The right graph depicts Wine Club ratings, with the horizontal axis ranging from ten to fifteen and the vertical axis showing relative percentages from zero to twenty percent. This graph similarly illustrates the frequency distribution, with a notable peak near fourteen ratings.Distributional comparison between the club’s review distribution and Platters’
Note(s): Histograms compare the distribution of Platter’s Guide expert ratings (1–5 stars) and Wine Club ratings (20-point scale). Platter ratings range from 1.5 to 5.0, with a mean of 3.92 and median of 4.0; Club ratings range from 6.0 to 18.6 (out of 20), with a mean of 15.1 and median of 15.3. The Platter distribution is left-skewed towards higher ratings, while the club distribution is more concentrated in the mid-to-upper range. These descriptive contrasts underscore the divergence in central tendency and spread between expert and consumer assessments
The image features two bar graphs side by side. The left graph presents Platter ratings on the horizontal axis, ranging from two to five, with the vertical axis displaying relative percentages from zero to thirty percent. The bars indicate the frequency of each rating, peaking at four. The right graph depicts Wine Club ratings, with the horizontal axis ranging from ten to fifteen and the vertical axis showing relative percentages from zero to twenty percent. This graph similarly illustrates the frequency distribution, with a notable peak near fourteen ratings.Distributional comparison between the club’s review distribution and Platters’
Note(s): Histograms compare the distribution of Platter’s Guide expert ratings (1–5 stars) and Wine Club ratings (20-point scale). Platter ratings range from 1.5 to 5.0, with a mean of 3.92 and median of 4.0; Club ratings range from 6.0 to 18.6 (out of 20), with a mean of 15.1 and median of 15.3. The Platter distribution is left-skewed towards higher ratings, while the club distribution is more concentrated in the mid-to-upper range. These descriptive contrasts underscore the divergence in central tendency and spread between expert and consumer assessments
It is relevant to note that some of the earliest scoring systems for wine used the 20-point scale. One notable example is the Davis scorecard [9]. This system emphasized deducting points for defects rather than rewarding positive qualities. The system only allocated two points for general quality. Although this method of scoring was effective in identifying faults, the Davis score card was less suited to expressing nuanced distinctions between high-quality wines (Lehrer, 2009)[10].
Research indicates that the ranking derived from the score average typically demonstrates greater accuracy compared to rankings based on the rank average or alternative methods like the Shapley ranking (Cao and Stokes, 2017). This suggests that using numerical scores directly may provide a more reliable reflection of wine quality compared to rank-based approaches because rank-based approaches can be influenced by ordinal scale limitations and ties.
In recent decades, the 100-point scoring system, introduced by Robert Parker, has become the standard for many wine critics and publications. Unlike the 20-point scales, Parker’s system starts at 50 and awards additional points for various attributes, such as colour, aroma, flavour and overall impression. Its popularity, particularly among American consumers, has been attributed to its resemblance to school grading systems. This has made it more intuitive and accessible (McCoy, 2005). The 100-point scale’s granularity allows for finer distinctions between wines, often influencing pricing and market demand significantly. However, critics of the system argue that it can oversimplify the complexity of wine, reducing subjective experiences to a single numerical value (Lehrer, 2009).
To assess the concordance between the Club’s and Platter’s wine ratings, a Cohen’s weighted kappa () was calculated. Weighted kappa adjusts for the severity of disagreements between categories (5 for Platter ratings and 20 for Club ratings, normalized to a common scale). It does this by assigning weights to different levels of disagreement. The weight is calculated as:
where n is the number of categories, is the absolute difference between the ratings assigned by the Club and Platter systems and determines the weight’s sensitivity to differences. For this analysis, was used to calculate the quadratic weighted kappa (QWK), which gives more weight to larger disagreements [11]. The weighted kappa statistic is then computed as:
where is the observed proportion of times the Club assigned rating i and Platter assigned rating j, and and are the marginal probabilities for the Club and Platter ratings, respectively.
The results indicated a very low level of agreement, with (95% CI: [0.003, 0.006]), based on 693 observations. The unweighted kappa, which does not consider the weight of disagreements, was 0, further underscoring the minimal concordance between the two rating systems. This finding aligns with the work of Hodgson (2008), who examined the reliability of wine judges in a major US competition and reported substantial inconsistency in scoring. His analysis revealed that individual judges often exhibited significant variability in their evaluations, with some applying narrow scoring ranges and others showing marked discrepancies even when assessing the same wines. Such variation, coupled with the low inter-judge agreement, raised concerns regarding the reliability of competition outcomes as indicators of intrinsic wine quality.
Our findings contribute by showing that both consumer−expert and expert−expert concordance in wine evaluation can be quite limited. Furthermore, studies that compare consumer and expert preferences directly (e.g. Goldstein et al., 2008) frequently observe substantial divergence, particularly for non-iconic wines or those in lower price tiers. Hodgson (2008)’s seminal findings in the US competition context are echoed in research from Australia (Gawel and Godden, 2008) and Europe (Ashton, 2011).
5. Expert price premiums and preferences
In this section, we leverage our unique longitudinal data set to examine how these different evaluations correlate with price premiums. Using a Bayesian framework, we model the relationship between CPI-deflated wine prices and ratings using a Gamma likelihood with a log link. Our motivation for using a Bayesian framework lies in the fact that this type of analysis offers several advantages over traditional frequentist approaches: non-Gaussian specification suited to strictly positive, right-skewed prices, direct estimation of full posterior uncertainty (with 95% credible/ highest posterior density [HPD] intervals) for parameters and derived quantities and the ability to make direct probabilistic statements about parameters and predictions [12]. The method also implements hierarchical (partial) pooling via random intercepts for yearbook and wine farm. This unique Bayesian structure shrinks noisy subgroup estimates towards the grand mean, stabilises inference in sparse cells and improves predictive performance, while retaining between group heterogeneity. Estimation is implemented in brms /Stan with standard diagnostics and posterior-predictive checks reported in Appendix 2 (Bürkner, 2017; Kruschke, 2021).
All wine prices were deflated to reflect 2024 prices using the consumer price inflation (CPI) index. To facilitate interpretation, all club scores are normalized to the same scale as Platter, with the minimum being cut off at 10. We rescale the Club’s 20-point scores to Platter’s five-point scale to enable direct comparability in coefficient magnitudes and graphical presentation later on in the paper [13]:
Building on Ali et al. (2008)’s findings, this analysis estimates the price premium associated with Platter’s ratings and evaluates whether the Club’s ratings also concur with the premium prices. By integrating expert and consumer perspectives, the study seeks to determine the extent to which these ratings converge or diverge in their impact on wine pricing. The following hypotheses are tested:
Higher Platter’s ratings are associated with higher deflated wine prices.
Higher club ratings are associated with higher deflated wine prices.
The relationship between ratings and wine prices varies by wine type.
To address the research questions and test the hypotheses, we use a Bayesian regression model with the following specification, where the independent variable deflated prices (y) is modelled by:
In this model, represents the global intercept term, capturing the baseline level of wine prices across all years. The vector of covariates includes main effects and interaction terms between Platter’s and club ratings and wine type. The associated coefficient vector reflects how ratings and their interactions with wine type relate to wine prices.
There might also be a temporal effect where the whole market for wine shifts due to unobservable factors. In contrast, simultaneity may also arise via (i) anticipatory pricing by producers expecting strong expert scores, (ii) extrinsic cue effects, where higher posted previous prices shape expert perception and (iii) omitted factors such as brand reputation and distribution that affect both ratings and price.
While we cannot measure the aforementioned factors directly, we model them as random deviations from the average price level using the yearbook edition. Likewise, unobserved heterogeneity between wine farms is accounted for through a second random intercept. These latent effects are captured through random deviations for yearbook and for wine farm. Both random intercepts are assumed to follow normal distributions, and .
The dependent variable, deflated wine price , is assumed to follow a Gamma distribution with mean and shape parameter . This distribution accommodates the positively skewed and non-negative nature of wine prices. The log link ensures that the conditional mean remains positive.
This formulation captures both the main effects of ratings on wine prices and their interaction with wine type, while also accounting for unobserved year- and wine farm-specific effects[14]. Estimation of the model was done using R and the brms package from Bürkner (2017).
5.1 Results
In this section, we analyse the results of the Bayesian model according to Kruschke (2021)BARG framework to elucidate the relationship between wine prices and the independent variables: Platter’s ratings, Club ratings. All model diagnostics are reported on in Appendix 2.
Table 2 and Figure 5 provide the estimates for the Bayesian model. The model provides evidence supporting the association between wine ratings and deflated prices, as well as insights into the variation in this relationship by wine type. All model estimates were back-transformed from the log scale to the original response scale for interpretability. Results are reported alongside 95% HPD intervals, which reflect highest posterior density credible intervals based on posterior draws.
Two graphs compare estimated wine prices based on ratings. Graph a labeled as Platter shows the relationship between Platter Rating, ranging from one to five on the horizontal axis, and Estimated Wine Price in Rand, measured on the vertical axis up to four hundred. Graph b labeled as Wine Club displays the Wine Club Rating, also ranging from one to five, with a vertical axis representing Estimated Wine Price in Rand, reaching up to four hundred as well. Both graphs include lines for red and white wines, represented by different colours. The data is arranged with the graph's lines showing trends across the ratings for both types of wine, highlighting changes in price associated with each rating category.Estimated marginal mean increases across price for Platter and Wine Club ratings
Note(s): Panels show estimated marginal means (EMMs) of deflated wine prices from the Bayesian regression. Shaded bands indicate 95% HPD intervals. Both expert (Platter’s Guide) and consumer (Club) ratings are positively associated with wine prices, with expert effects stronger in magnitude. For red wines, predicted prices rise from 89.8 (95% HPD: [64.9, 118.0]) at a Platter score of 1 to 296.0 ([242.0, 348.0]) at a score of 5, while Club ratings correspond to increases from 112.0 ([83.9, 143.0]) to 237.0 ([186.0, 293.0]). White wines follow a similar pattern but at consistently lower levels: Platter ratings increase from 59.5 ([42.8, 77.0]) to 218.9 ([179.0, 259.0]), while Club ratings increase from 86.2 ([61.2, 110.0]) to 150.0 ([113.0, 192.0])
Two graphs compare estimated wine prices based on ratings. Graph a labeled as Platter shows the relationship between Platter Rating, ranging from one to five on the horizontal axis, and Estimated Wine Price in Rand, measured on the vertical axis up to four hundred. Graph b labeled as Wine Club displays the Wine Club Rating, also ranging from one to five, with a vertical axis representing Estimated Wine Price in Rand, reaching up to four hundred as well. Both graphs include lines for red and white wines, represented by different colours. The data is arranged with the graph's lines showing trends across the ratings for both types of wine, highlighting changes in price associated with each rating category.Estimated marginal mean increases across price for Platter and Wine Club ratings
Note(s): Panels show estimated marginal means (EMMs) of deflated wine prices from the Bayesian regression. Shaded bands indicate 95% HPD intervals. Both expert (Platter’s Guide) and consumer (Club) ratings are positively associated with wine prices, with expert effects stronger in magnitude. For red wines, predicted prices rise from 89.8 (95% HPD: [64.9, 118.0]) at a Platter score of 1 to 296.0 ([242.0, 348.0]) at a score of 5, while Club ratings correspond to increases from 112.0 ([83.9, 143.0]) to 237.0 ([186.0, 293.0]). White wines follow a similar pattern but at consistently lower levels: Platter ratings increase from 59.5 ([42.8, 77.0]) to 218.9 ([179.0, 259.0]), while Club ratings increase from 86.2 ([61.2, 110.0]) to 150.0 ([113.0, 192.0])
Summary of Bayesian model estimates
| Parameter | Estimate | Std. error | 2.5% CI | 97.5% CI | Rhat | ESS |
|---|---|---|---|---|---|---|
| Fixed effects | ||||||
| Intercept | 38.1 | – | 23.3 | 61.6 | 1.00 | 1207 |
| Platter score | 1.35 | – | 1.22 | 1.48 | 1.00 | 1376 |
| White wine | 0.75 | – | 0.40 | 1.38 | 1.00 | 1435 |
| Club score | 1.21 | – | 1.09 | 1.32 | 1.00 | 1603 |
| Platter score white wine | 1.03 | – | 0.91 | 1.16 | 1.00 | 1541 |
| Club score white wine | 0.95 | – | 0.83 | 1.08 | 1.00 | 1501 |
| Random effects | ||||||
| Wine farm | ||||||
| sd(intercept) | 1.31 | 0.03 | 1.23 | 1.39 | 1.00 | 803 |
| Yearbook | ||||||
| sd(intercept) | 1.45 | 0.05 | 1.32 | 1.62 | 1.01 | 510 |
| Distributional parameter | ||||||
| Shape | 5.91 | 0.40 | 5.15 | 6.70 | 1.00 | 1057 |
| Parameter | Estimate | Std. error | 2.5% | 97.5% | Rhat | |
|---|---|---|---|---|---|---|
| Fixed effects | ||||||
| Intercept | 38.1 | – | 23.3 | 61.6 | 1.00 | 1207 |
| Platter score | 1.35 | – | 1.22 | 1.48 | 1.00 | 1376 |
| White wine | 0.75 | – | 0.40 | 1.38 | 1.00 | 1435 |
| Club score | 1.21 | – | 1.09 | 1.32 | 1.00 | 1603 |
| Platter score | 1.03 | – | 0.91 | 1.16 | 1.00 | 1541 |
| Club score | 0.95 | – | 0.83 | 1.08 | 1.00 | 1501 |
| Random effects | ||||||
| Wine farm | ||||||
| sd(intercept) | 1.31 | 0.03 | 1.23 | 1.39 | 1.00 | 803 |
| Yearbook | ||||||
| sd(intercept) | 1.45 | 0.05 | 1.32 | 1.62 | 1.01 | 510 |
| Distributional parameter | ||||||
| Shape | 5.91 | 0.40 | 5.15 | 6.70 | 1.00 | 1057 |
The model uses a Gamma family with a log link for the mean and an identity link for the shape. The data set contains 693 observations, and posterior distributions were sampled using four chains with 1,000 iterations each (500 warmup), resulting in 2,000 total post-warmup draws. All reported fixed-effect estimates and their credible intervals have been transformed to the response scale using , given the log-link specification. Random-effect standard deviations are shown as multiplicative factors on the response scale. The shape parameter is reported on its original scale. Std. errors are not shown as the column is a standard error on the log scale and because transformation is non-linear, and taking the would not be representative of the actual standard error
The intercept () (38.1; 95% CI: [23.3, 61.1]) represents the baseline deflated price for red wines when Platter’s and Club ratings are at their reference levels. Both Platter’s and Club ratings were positively associated with deflated prices, supporting the hypotheses that higher ratings from experts and consumers are associated with increased prices. A one-unit increase in Platter’s ratings was associated with a 35% increase in price (1.35; 95% CI: [1.22, 1.48]). This strong and significant association underscores the prominent role of expert evaluations in shaping wine prices. Similarly, a one-unit increase in club ratings was associated with a 21% increase in deflated price (1.21; 95% CI: [1.09, 1.32]), highlighting the relevance of consumer preferences in association with wine prices, though with a smaller effect size compared to expert ratings.
For Platter’s ratings, the interaction term (1.03; 95% CI: [0.91, 1.16]) indicated a negligible attenuation of the price impact for white wines. However, the credible interval includes 1, suggesting no significant difference in the effect of expert ratings between red and white wines. Similarly, the interaction between club ratings and wine type (0.95; 95% CI: [0.83, 1.08]) suggested a modest reduction in the price impact for white wines. While the effect was not statistically meaningful, as the credible interval includes 1, this finding suggests consistency in the evaluation of wine types, both for consumers and experts alike [15].
The estimated marginal means (EMMs) reveal that both consumer (Wine Club) and expert (Platter’s) ratings are significantly associated with wine prices, with expert ratings exerting a stronger effect. For red wines rated by the club, predicted prices increase steadily with higher ratings, ranging from 112.0 ([83.9, 143.0]) at a score of 1.0 to 237.0 ([186, 293.0]) at a score of 5.0. Similarly, for Platter’s ratings, red wine prices rise from 89.8 ([64.9, 118.0]) at a score of 1.0 to 296.0 ([242.0, 348.0]) at a score of 5.0. White wines exhibit comparable trends, with prices increasing with higher ratings but consistently remaining below those of red wines for equivalent scores. For example, white wine prices increase from 86.2 ([61.2, 110.0]) to 150.0 ([113.0, 192.0]) across the range of club ratings, while prices for Platter’s ratings increase from 59.5 ([42.8, 77.0]) to 218.9 ([179.0, 259.0]) over the same score range.
Finally, estimated random effects indicate that both producer- and time-specific factors shape wine prices. Variation across wine farms corresponds to a multiplicative factor of 1.31 ([1.23, 1.39]), meaning that a wine with a baseline predicted price (based on rating) of R150 could range from approximately 115 to 195 simply due to differences between producers. Taken together, these results highlight that both farm-level heterogeneity and temporal fluctuations contribute substantially to observed pricing patterns.
In summary, the results provide robust evidence supporting H1 and H2: higher Platter’s and Club ratings are significantly associated with higher price premiums. Specifically, Platter’s ratings exert a stronger influence on price premiums compared to club ratings. This is consistent with prior findings emphasizing the market power of expert evaluations (Ali et al., 2008; Lecocq and Visser, 2006; Ashton, 2016). A one-unit increase in Platter’s ratings corresponds to an approximate 35% increase in price, whereas the equivalent increase for club ratings is 21%.
In contrast, H3, positing differential effects across wine types, finds no significant empirical support. The interaction terms between ratings and wine type exhibit overlapping credible intervals, suggesting that the effect of ratings on price does not meaningfully vary between red and white wines. Nevertheless, a modest difference is observed in the price trajectories: for the club ratings, the price premium for white wines grows more slowly than for reds, while for Platter’s ratings, the effect size remains proportional across wine types, albeit from a lower base for white wines.
These findings reinforce the substantial role of both expert and consumer evaluations in wine pricing mechanisms. In particular, the dominance of expert ratings aligns with prior meta-analyses highlighting the premium attached to authoritative reviews in hedonic pricing models (Oczkowski and Doucouliagos, 2015; Villas-Boas et al., 2021; Bazen et al., 2024). Additionally, the consistency of the rating−price relationship across wine types suggests that market participants do not systematically differentiate between reds and whites in their responsiveness to quality signals. The latter contributing to the broader literature on price formation in differentiated product markets (Masset and Raub, 2023; Schiefer and Fischer, 2008).
From a practical perspective, these results suggest that wineries and marketers should continue to emphasize expert endorsements in their positioning strategies, given their stronger price impacts relative to consumer-driven ratings. Additionally, the stable rating effects across varietals imply that the concept of a premium wine can be applied across both red and white wine categories without significant adjustment for type. It should be noted, however, that our analysis is confined to retail list prices. In practice, price premia attached to ratings are likely to vary substantially across channels (e.g. restaurants, online discounting or “mystery wine” offerings), and the magnitude of a point difference at the very top of the scale may be attenuated dependent on the settings.
5.2 Discussion
While most studies cited in this paper are based in the USA or Europe, it is worth reflecting on how these results relate to the South African market context. The consistent price premiums observed, despite the low concordance between expert and consumer ratings, may be partly explained by structural and cultural features of the domestic wine industry. South Africa’s premium wine market remains relatively concentrated, with a limited number of producers capturing most of the attention from critics, distributors and consumers [16]. This fosters conditions where expert scores can exert outsized influence, especially among higher-income or export-oriented segments.
This reliance of consumers in South African on experts may place heightened value on their authority due to the country’s historical reliance on a small number of gatekeepers in wine journalism and retail (Priilaid, 2006). Experimental and neuroeconomic work also shows that consumers respond strongly to extrinsic cues such as medals and expert scores, often over and above intrinsic quality (Priilaid and van Rensburg, 2012; Plassmann et al., 2008; Masset and Raub, 2023).
Our findings indicate that expert and consumer ratings function as complementary (partially orthogonal) quality signals. Even with low agreement between Platter’s Guide and the Club, both signals are positively associated with price (with the expert effect stronger in magnitude). This has two implications. Firstly, it suggests audience segmentation: more expert‐attentive segments (premium/collectible/export) are likely to respond primarily to critic scores, whereas socially influenced segments (retail/online) may respond more to community ratings (Babin and Bushardt, 2019; Kaimann et al., 2023; Bazen et al., 2024). Secondly, it implies cue complementarities: displaying both consumer and expert ratings can raise perceived quality. Expert endorsements serve as a signal of authority, while community ratings as social proof (Friberg and Grönqvist, 2012; Villas-Boas et al., 2021).
From a producer and intermediary perspective, the aim would be to align information cues with the characteristics of the distribution channel and marketer. In premium and direct‐to‐consumer context (where search costs for expert knowledge are lower and reputational capital is salient), expert evaluations should be displayed more prominently. In mass retail and online environments, where social information is comparatively more diagnostic, a combination of expert scores with community (consumer) ratings may generate complementary signalling effects (authority and social proof). In markets where brand reputation is already strong (Oczkowski, 2018), ratings can be used to calibrate amongst peers and price points rather than determine overall positioning.
6. Conclusion
This study provides new empirical evidence on the relationship between wine ratings and price formation. The paper leverages a rare longitudinal data set spanning nearly four decades from a private Stellenbosch-based wine tasting club. The analysis integrates expert evaluations from Platter’s South African Wine Guide with aggregated consumer ratings, offering a dual perspective on quality assessment and the formation of price premiums in the wine market.
Despite the Club’s composition of knowledgeable and experienced tasters, agreement with expert ratings (Platter's) was notably low. The near-zero weighted kappa statistic indicates minimal concordance between consumer and expert assessments. This finding strengthens the known subjectivity and variability inherent in wine evaluation (Hodgson, 2008; Schiefer and Fischer, 2008; Goldstein et al., 2008). The low agreement suggests that, even among informed consumers, alternative factors shape wine preferences in ways that diverge from expert norms and intrinsic quality metrics. Despite low agreement between the expert and consumer ratings, low , both independently predict wine prices. This highlights a key insight: Platter’s and the Club’s ratings capture distinct, partially orthogonal dimensions of perceived quality.
A Bayesian price model confirms that higher ratings from both Platter’s (H1) and the Club (H2) are associated with significant price premiums. The stronger association for expert ratings aligns with research on the signalling value of professional evaluations (Ali et al., 2008; Lecocq and Visser, 2006; Ashton, 2016; Villas-Boas et al., 2021) and extends earlier South African findings on non-linear rating−price relationships (Priilaid and van Rensburg, 2012). H3, proposing differential effects by wine type, finds limited support: although red wines command higher baseline prices, the marginal effects of ratings are largely consistent across red and white wines.
Taken together, these findings contribute to the wine business literature by illustrating how expert and consumer evaluations coexist and influence market outcomes, even when they diverge observationally. Practically, the results suggest that wineries and marketers could benefit from presenting both expert and consumer ratings. This would be particularly true in heterogeneous markets where salient signals drive consumer preferences.
Notes
At the opposite end, “alcohol <13%” (), “region of origin” (), “information on shelf”' (), “read about it” () and “information on back label” () were least important. This suggests that they have limited reliance on information-rich, commercial sources. Price was not included in the 13-attribute BWS design, although a small number of respondents () volunteered it as important.
The advantage of following mostly the same producers and wine types (with the same cohort) over such a long period provides within-producer and within-year variation that sharpens identification of rating–price associations. At the same time one needs to be aware that long panels are often imbalanced with producer entry/exit or occasional missing scores/prices. They are also susceptible to measurement drift (e.g. evolving evaluation practices or club composition).
Bruwer et al. (2011) report specific differences in wine consumption behaviour and sensory preferences by gender and generation. Females tend to consume less wine overall and spend less in total, but pay more per bottle (possibly as a risk‐reduction strategy). Female consumers show higher consumption of white wines, a preference for sweeter styles at a younger age and stronger preferences for “medium‐bodied wines with fruit‐forward, vegetative, oak and mouth‐feel characters”. Males more often prefer aged wine characteristics. Other related work links gender to differences in information use and packaging/label sensitivity when selecting wine, and cues such as cap or cork can influence perception (Barber et al., 2006).
See Appendix 1 for a disaggregated view per wine type.
The reason why December is red wine dominant is due to low tasting occurrences during this holiday month: red (10) and white (6). The chi-squared results indicate a significant association between wine type and month ( = 202.69, df = 11, p < 2.2 × 1016). The results show that the distribution of preferences for red and white wines is not uniform across months, implying a temporal component to consumer wine preferences in South Africa.
The estimation uses a linear regression model where the dependent variable is cellaring time and the independent variable is year (cellaring year). The model yields a residual standard error of 1.746 on 329 degrees of freedom, with an adjusted of 0.2195. The F-statistic of 4.085 on 33 and 329 degrees of freedom is statistically significant (p = 1.672 10-11), indicating that the model explains a significant portion of the variation in cellaring time.
Involves calculating the simple arithmetic mean of numerical scores assigned by judges.
Based on averaging the ranks assigned to wines by judges.
Originating from the University of California, Davis.
Another influential 20-point system is the Roseworthy scorecard that was developed in Australia at the Roseworthy Agricultural College. This system aimed to refine wine assessment by combining objective criteria with sensory evaluation. The idea behind this score card was to help formalize wine judging in the Southern Hemisphere.
The quadratic weighted kappa (QWK) is often preferred for ordinal scales.
For example, .
As a robustness check, the Club’s ratings were kept on the original 20-point score card. Although the coefficients adjust slightly, the results and inference remain the same.
By incorporating the random intercept, the model accounts for hierarchical structure, providing more robust estimates of the fixed effects. As a robustness, the year and wine farm random variables were left out and the expected log predictive density and effective number of parameters in the Leave-One-Out Information Criterion (LOOIC) metrics indicated the model fit to favour the inclusion.
Switching from red to white wine was associated with a 25% decrease in price (0.75; 95% CI: [0.48, 1.38]). However, the wide credible interval indicates substantial uncertainty. This suggests that differences in baseline prices between red and white wines vary considerably across the sample.
Although there has been an increase in the number of wineries (cellars) in South Africa increased from under 300 in 1997 to almost 600 in 2010, many of these are small operations. The proportion of cellars crushing less than 100 tonnes of grapes annually rose from one-quarter of the total in 1997 to nearly half in recent years (Vink et al., 2012). In a study by SAWIS and Frost & Sullivan/FTI Consulting (2022), they show that the number of primary grape producers has fallen by a quarter over the last decade, from 3,323 in 2013 to 2,487 in 2022, with most producing less than 500 tonnes of grapes per year.
References
Appendix. Wine club ratings per year
Figure A1 provides an overview of how many of each wines type were tasted each year. It highlights the differing patterns in tasting frequency for red, rosé, sparkling and white wines across five year intervals.
This bar graph illustrates the count of wines per year between 1990 and 2020, divided into four categories: red, rosé, sparkling, and white. The vertical axis represents the number of wines, ranging from zero to twenty. The horizontal axis shows the years in five-year increments. The graph contains four distinct rows for each wine type, labelled at the top of each row. The red and white wine categories show considerable variations in counts over the years, with multiple peaks, while the rosé and sparkling categories have fewer instances with sparse tallies in certain years. Each category is visually differentiated using distinct colours: red wines appear in dark maroon, rosé in orange, sparkling in dark blue, and white in light blue. The layout flows vertically, arranged by wine type, making it straightforward to compare the counts across the specified years.Number of wines tasted per year
This bar graph illustrates the count of wines per year between 1990 and 2020, divided into four categories: red, rosé, sparkling, and white. The vertical axis represents the number of wines, ranging from zero to twenty. The horizontal axis shows the years in five-year increments. The graph contains four distinct rows for each wine type, labelled at the top of each row. The red and white wine categories show considerable variations in counts over the years, with multiple peaks, while the rosé and sparkling categories have fewer instances with sparse tallies in certain years. Each category is visually differentiated using distinct colours: red wines appear in dark maroon, rosé in orange, sparkling in dark blue, and white in light blue. The layout flows vertically, arranged by wine type, making it straightforward to compare the counts across the specified years.Number of wines tasted per year
Bayesian model diagnostics
Traceplots
Figure A2 displays traceplots for all estimated parameters across the four MCMC chains. The chains exhibit good mixing and appear stationary after the warmup phase, with no signs of divergence or poor exploration of the posterior space. Visual inspection confirms convergence for all parameters. One minor deviation was observed in Chain 3 for a single coefficient, which shows a slight shift in density mass relative to the others, though this deviation is negligible and does not materially affect posterior summaries.
The image presents four line graphs arranged in a 2x2 grid. Each graph represents one parameter: b_Intercept, b_average_score_platters, b_wine_club_norm, and b_typewhite, displayed in order from top left to bottom right. The x-axis spans values from 500 to 1000, indicating a range of iterations or samples. The y-axes vary by parameter but show changes within their respective ranges, with b_wine_club_norm mainly confined to values between 0.1 and 0.4. Each line within the graphs corresponds to a different statistical chain, indicated by a key in the lower left corner, showing four chains plotted in different colours.Traceplots for all model parameters across four chains
The image presents four line graphs arranged in a 2x2 grid. Each graph represents one parameter: b_Intercept, b_average_score_platters, b_wine_club_norm, and b_typewhite, displayed in order from top left to bottom right. The x-axis spans values from 500 to 1000, indicating a range of iterations or samples. The y-axes vary by parameter but show changes within their respective ranges, with b_wine_club_norm mainly confined to values between 0.1 and 0.4. Each line within the graphs corresponds to a different statistical chain, indicated by a key in the lower left corner, showing four chains plotted in different colours.Traceplots for all model parameters across four chains
Posterior density plots
Posterior density plots for all model parameters are shown in Figure A3. The densities appear smooth and unimodal, with no indication of multimodality or erratic behaviour, suggesting well-behaved posterior distributions. Convergence diagnostics further support these visual impressions. One commonly used diagnostic is the Gelman–Rubin statistic (), which assesses whether the independent Markov chains have converged to the same target distribution. Values of close to 1.00 indicate good convergence; all parameters had ≤ 1.01, with most exactly equal to 1.00.
Another key diagnostic is the effective sample size (ESS), which measures the number of independent samples equivalent to the correlated MCMC draws. It is reported for both the bulk of the posterior distribution (Bulk ESS) and its tails (Tail ESS). Larger ESS values imply more reliable estimation. In this model, all Bulk ESS values for regression coefficients exceeded 900 − for example, 1,336 for Platter Score and 1,284 for Wine Club − indicating high precision in posterior summaries. Tail ESS values were similarly robust, all exceeding 1,200. The standard deviation of the year-level random intercept, which draws on fewer observations, still achieved acceptable values (Bulk ESS = 310; Tail ESS = 727).
Together, these diagnostics confirm that the model chains mixed well, converged successfully, and generated a sufficient number of effective samples to support robust posterior inference.
The image contains four separate graphs arranged in a two-by-two layout. Each graph presents the distribution of a different variable, labelled at the top left of each plot: b_Intercept in the top left, b_average_score_platters in the top right, b_typewhite in the bottom left, and b_wine_club_norm in the bottom right. Each graph features lines representing four chains, denoted by different colours. The horizontal axis represents the variable's values, while the vertical axis indicates the density. For the graphs with the variables b_Intercept and b_average_score_platters, the scale on the vertical axis ranges from zero to a maximum of approximately eight, while the b_typewhite and b_wine_club_norm graphs show similar density ranges but start at zero and appear lower in scale. Each graph includes a grid for reference and a legend at the bottom for identifying the chain representations, with labels for Chains 1, 2, 3, and 4.Posterior density estimates for all model parameters
The image contains four separate graphs arranged in a two-by-two layout. Each graph presents the distribution of a different variable, labelled at the top left of each plot: b_Intercept in the top left, b_average_score_platters in the top right, b_typewhite in the bottom left, and b_wine_club_norm in the bottom right. Each graph features lines representing four chains, denoted by different colours. The horizontal axis represents the variable's values, while the vertical axis indicates the density. For the graphs with the variables b_Intercept and b_average_score_platters, the scale on the vertical axis ranges from zero to a maximum of approximately eight, while the b_typewhite and b_wine_club_norm graphs show similar density ranges but start at zero and appear lower in scale. Each graph includes a grid for reference and a legend at the bottom for identifying the chain representations, with labels for Chains 1, 2, 3, and 4.Posterior density estimates for all model parameters
The image presents a graph that shows two curves, one labeled Y and the other Y rep, plotted against a horizontal axis ranging from zero to two thousand. The vertical axis represents values that range from zero to approximately zero point zero three. The curves demonstrate a decreasing trend as they move rightward across the axis, indicating a systematic decline. The graph uses thin, overlapping lines to illustrate the data points, creating a cloud of curves that visually represent the values for both Y and Y rep over the specified range.Posterior predictive distribution (dark) overlaid on observed data (light)
The image presents a graph that shows two curves, one labeled Y and the other Y rep, plotted against a horizontal axis ranging from zero to two thousand. The vertical axis represents values that range from zero to approximately zero point zero three. The curves demonstrate a decreasing trend as they move rightward across the axis, indicating a systematic decline. The graph uses thin, overlapping lines to illustrate the data points, creating a cloud of curves that visually represent the values for both Y and Y rep over the specified range.Posterior predictive distribution (dark) overlaid on observed data (light)
Posterior predictive check
To evaluate the overall fit of the model, we conducted a posterior predictive check, a standard diagnostic procedure in Bayesian analysis. This method involves simulating new data from the posterior distribution of the model and comparing it to the observed data. The idea is straightforward: if the model is a good fit, the simulated (predicted) data should resemble the actual data.
Figure A4 displays the posterior predictive distribution (or PP Check) overlaid on the observed distribution of the deflated wine prices. Visually we can see the model captures the pronounced right-skewed shape of the empirical distribution, which is characteristic of price data and consistent with our choice of a Gamma likelihood combined with a log-link function. Importantly, the predictive distribution aligns well with the observed data across the entire range, including the long right tail, where price outliers often reside.
This result provides strong evidence that the model is appropriately specified for the underlying data generating process. The close correspondence between observed and predicted distributions suggests that key structural features of the data − such as skewness, scale and spread − are accurately captured by the model. As such, the posterior predictive check supports the adequacy of the model for inference purposes.

