Will AI Replace Wine Criticism?

by Dwight Furrow

In World of Fine Wine, Master of Wine Benjamin Lewin speculates about whether AI could replace human-generated wine criticism. He acknowledges the obvious limitation: “A chatbot cannot smell or taste.” Yet he thinks this may not matter. Wine, after all, consists of measurable properties. The intensity of color and its wavelength can be specified. Labs can analyze acidity, sugar, alcohol, viscosity, and volatile compounds with far more precision than the human palate. Feed those measurements into an AI trained on the enormous archive of published tasting notes and scores, Lewin argues, and the machine could spit out the patterns linking chemical composition to human judgment about taste. Impressions would be replaced by actual measurement, he suggests. Eventually, what he calls “a self-driving laboratory” (analogous to a self-driving car) might generate tasting notes, predict aging, adopt different critical “personas,” and achieve “greater objectivity than any single human expert.”

Common sense might suggest that the fact that an LLM cannot smell or taste should end the discussion. But that limitation does not harm Lewin’s argument. His point is that an LLM, supplied with laboratory data along with the vast archive of online tasting notes, could accurately predict what informed tasters would find in a wine. An AI that can reliably predict what expert tasters would report would be functionally equivalent to what actual tasters discover.

I do not think we can dismiss this speculation out of hand. Machines may become very good at predicting sensory reports. They could easily track chemical changes across thousands of wines or compare a new wine with thousands of prior examples. But whether we should take this hypothesis seriously depends in part on what we think the purpose of wine criticism is. If criticism is primarily about classifying wines according to established descriptors and quality rankings, then Lewin’s proposal is plausible. If criticism should make a wine intelligible as an aesthetic object, the case becomes more difficult.

Measurement is not perception. I am sure Lewin knows this, but I’m not sure he gives it sufficient attention. Sugar concentration differs from perceived sweetness, and pH doesn’t tell you how the acidity feels. A mass spectrometer measures volatile compounds in a wine, but their sensory effects depend on the wine’s structure, oxygen exposure, the interactions among compounds, and the physiology and attention of the taster. But more importantly, the way aromatic compounds interact with human sensory mechanisms is enormously complex and non-linear, and so doesn’t lend itself to straightforward predictions. None of this proves that AI cannot model perception. A machine needn’t smell in order to predict what human beings will smell, any more than a medical information system would have to feel pain to estimate who is likely to suffer it. But the inference from chemistry to perception is far more complicated than Levin’s hypothesis seems to allow.

Lewin also moves too quickly from description and recommendation to evaluation. A machine may measure compounds in a wine, recommend bottles to a consumer, predict the language critics will use, and estimate the score a panel will assign. But is this providing an evaluation of the wine? An AI system trained on historical scores learns that certain combinations of ripeness, oak, alcohol, acidity, and texture tend to receive high ratings. That tells us what people like. But it does not tell us why those wines deserve approval. That is a different sort of question.

Much wine criticism is, indeed, sort of mechanical, consisting of aroma descriptors, judgments about what is typical, cellar recommendations, and scores. An AI may mimic that style, but informed criticism should explain how a wine produces its effects. Wine unfolds. Aromas rise, fade, and return in altered form. Acidity may initially sharpen the fruit, then expose some bitterness. The wine might initially seem tight and unexpressive but open up with aeration. The finish may resolve the wine’s tensions or send the wine somewhere unexpected. A critic tracks this movement. Does the fruit gain definition from the acidity, or does the acidity simply sit beside it? Does volatile acidity animate the wine or cause its structure to fall apart? These are questions about the wine’s organization in time and how expressive the wine is, and good criticism should give us some of that. And some wine critics do. Their descriptive language is controversial and largely metaphorical, but it can give readers a sense of a wine’s expressiveness. A wine may appear severe and then become generous. It may be expansive at first and then become increasingly austere, nervously energetic, languid, compressed, or playful. The critic’s job is to articulate such patterns so that another taster can notice them and test them against their own judgment.

Could an AI learn these temporal and expressive relations from chemical analysis? Possibly. I suppose dynamic measurements could track the release of volatile aromas and track chemical developments over time. Human reports could supply data about attack, mid-palate, finish, and aromatic changes as the wine sits in the glass. Perhaps the metaphors can be cashed out in more literal language. I am not suggesting there is some mystical barrier between machines and meaning. But Lewin’s hypothesis largely ignores the problem. His imagined system begins with analytical data and ends with a tasting note, while the organization of sensation, the part that makes the tasting note criticism rather than a list of features, is left out.

Finally, Lewin treats subjectivity chiefly as noise. Critics are inconsistent, he points out. Their taste and moods change from day to day. They are vulnerable to changes in fashion, give undue weight to reputation or give in to commercial pressure. Perhaps AI would be less susceptible to these biases.

Yet there is a form of bias that we don’t want to eliminate. We want an assessment generated by the critic’s informed point of view. A critic’s perspective is a cultivated capacity to notice features of a wine that develops through the practice of constantly comparing wines, arguing about them, calibrating their judgements with the views of other experts, and revising their views over time. We want judgements to be made from that informed perspective because perspectives give us useful information. One critic may be especially alert to structure, another may focus on aromatic development, and another might be interested in the way a wine departs from regional conventions. These differences give us access to features of the wine that a supposedly neutral summary might not pick up.

Lewin’s suggestion that an AI could adopt different critical “personas” is therefore more revealing than he intends. If the machine can speak as a classicist, a natural-wine advocate, or a partisan of power and ripeness, it will have learned to simulate several well-established contemporary ideologies. But this does not add up to a point of view of its own shaped by a history of encounters with particular wines that allow it to look beyond current trends. A good critic’s preferences will not simply fall into a well-defined category. We are often surprised by what we like or don’t like. Human preferences are not consistent nor should they be. That violation of expectation makes wine tasting exciting especially when the variation is meaningful. Perhaps an AI can be “surprised”—it can certainly detect unusual patterns in a body of data. But whether that “surprise” is meaningful or not? That requires human judgment.

The role of surprise and the violation of expectations is important because wine appreciation, and critical judgment at its best, involves the pursuit of wines that are singular. A wine becomes singular when its individuality resists easy substitution: when it cannot be understood as merely another example of its grape, region, style, or score category. Such wines put pressure on the categories used to judge them. A good critic goes beyond registering deviation from a norm to explain why the deviation matters, at least to her, from her perspective.

It is not obvious how an LLM discovers that. An AI might detect anomalies in both chemistry and language. But an anomaly is not a singularity. An anomaly is statistically unusual. A singular achievement is a difference whose importance has to be interpreted. And so criticism sometimes requires invention. The critic confronts a wine for which the available vocabulary is inadequate and must find a new description, perhaps even revising the standard by which the wine is judged. A model trained on the history of criticism may reproduce the judgments through which earlier singular wines became intelligible. Lewin gives us little reason to think it can recognize the next one before human critics have named it and described it to us.

AI will almost certainly become useful to wine criticism. It may expose inconsistencies in wine discourse and improve recommendations because of the sheer amount of data to which it has access. It may even force critics to become less lazy, which would be no small public service. But Levin has not shown that such a system can replace criticism in the fullest sense. Maybe some future AI acquires the kind of judgment required of wine (or art, literature, or music) criticism but before declaring the critic obsolete, we should be very clear about what criticism is supposed to accomplish.