How Facts Are Made

by W. Alex Foxworthy

Recently, following an assignment that required AI use, a student asked me whether it was true that every question she asks a chatbot uses up a whole bottle of water. Like so much information today, she received this claim via a social media video. The video apparently cited a newspaper article which cited a real study. She had also heard people all around her, both in real life and online, talking about the huge amounts of water these systems use. Of course, she could get more information by searching online, asking a chatbot, asking friends and teachers or listening to the many podcasts discussing the issue. In other words, my student has no shortage of information, but what she doesn’t have is a solid way to judge its accuracy, how much confidence a given claim deserves, nor which sources can be trusted. Though many students are now “digital natives,” in the sense that they have grown up with smartphones and tablets, it has been my experience that they are not skilled in navigating the bewildering and fast growing world of information available to us all. 

At the same time, there seems to be a deepening distrust in institutions and authorities as well as widespread disagreement on the factual basis of some of the most important issues in the public discourse. More than ever, I hear from students who don’t see the value of a “well-rounded” education nor in spending time in courses that aren’t directly tied to their career paths. Artificial intelligence has made the question of what an education offers even more urgent, as we now have machines that deliver confident answers on almost any subject (some right and some wrong) and produce content far faster than people can check it. Taken together, these observations have forced me to think about what my role in this educational process is really for. Particularly when it comes to teaching non-science majors, I think what science has to offer are ways of thinking that can be applied across many domains. The framework key to it all is something we now describe with the shorthand “scientific method”: how human beings have learned to test claims about the world and preserve what survives rigorous testing. Of course, most science courses already cover the scientific method: ask a question, form a hypothesis, test it, analyze the results, and draw a conclusion. But the treatment is often superficial, students don’t arrive in my classroom with an intuitive understanding of what it really means, and I think we leave out much of what makes the process work. This essay is my attempt to lay out what the five steps leave out, using my student’s question as an illustrative example.

Most of us still judge claims using capacities that long predate any formalized scientific process. Instinct provides responses shaped by natural selection over many generations. Newborn babies turn towards the breast without being taught, and we withdraw rapidly from painful stimuli. These responses help us survive, but they do not necessarily give us a deep or unbiased understanding of the world. Individual experience adds a faster layer of learning. However, it can sometimes be misleading as we tend to credit whatever preceded a good outcome, whether or not it actually caused it. Experience also has difficulty revealing a cause that acts slowly over years or only changes the probability of an outcome. A third method we rely on is tradition: what “everyone” knows and what your people always do. Through oral traditions, written language, and other durable forms of communication, communities can preserve knowledge beyond what any individual could work out. However, they can hold on to errors quite securely as well. Often a traditional practice and its explanation come bundled together, and an apparently successful practice lends credibility to an untested explanation. Bloodletting persisted for centuries despite challenges to its usefulness. A patient’s recovery after being bled could seem to confirm the treatment, even when the patient would have recovered without it. A fourth way of knowing is argument using reason and logic. These methods allow people to examine ideas and argue for or against them in ways that inherited practices may resist. However, just because something makes sense doesn’t mean that it’s true. Even valid reasoning cannot establish that its starting premises are correct. Even today, science continues to rely on experience, reasoning, and knowledge passed down by others. Much of the scientific process comes down to how we check our knowledge.

The typical science textbook’s answer to this is the experiment. Students are taught to come up with an idea, test it against the world, and use the results to refute or support it. However, this sounds like nothing more than checking your assumptions, which raises an obvious question. Haven’t people always checked their assumptions? An ancient carpenter might think a plank was long enough to span an opening, then measure both and discover that it was too short. He could even lay the plank across the opening and see for himself. If experimentation means only comparing an expectation with reality, it is hard to imagine that people waited until the seventeenth century to try it. What I think is missing from this picture is that ordinary observation quite often supports an incorrect explanation. For example, Aristotle held that heavier bodies fall faster in proportion to their weight. In fact, if you drop a stone and a leaf, the result seems to support the general idea that heavier objects fall faster. Watching the sun cross the sky similarly gives the impression that it is moving around a stationary Earth. The world as it presents itself to naive observation is often misleading.

The scientific process doesn’t just hold that claims should be checked. Instead, the check must be carefully arranged so that we have a meaningful chance of finding the claim wrong. Sometimes we build the comparison by removing what confounds the question: for example, removing the air so that air resistance doesn’t obscure the relationship between weight and falling motion. Other times, investigators find comparisons in the natural world that allow them to distinguish between competing explanations. Darwin tested explanations of organismal descent, variation and geographic spread through careful observations in the natural environment, alongside his experimental work. As Darwin stated in a letter, “all observation must be for or against some view if it is to be of any service.” The important thing is not simply to look carefully at nature, but to seek observations for which competing explanations predict different results.

The investigator’s own expectations can also interfere with a comparison. An influential early example of a precaution against this problem was the British trial of streptomycin for tuberculosis, reported in 1948. Treatment assignments followed a sequence based on random numbers, concealed in sealed envelopes. Investigators could not know a patient’s assignment before the patient was accepted into the trial. Even a random sequence could have been undermined if doctors knew the next assignment and used that knowledge to decide whom to enroll. Concealing it helped prevent them from consciously or unconsciously steering particular patients toward one group. The procedure protected the comparison from the preferences of the people conducting it, including doctors who understandably hoped to help their patients.

These precautions developed over centuries, but carefully arranged comparisons existed long before the seventeenth century. In sixth-century Alexandria, John Philoponus described a straightforward test of Aristotle’s account of falling bodies: drop two objects of very different weights from the same height. If they fell faster in proportion to their weights, an object twice as heavy should cover the distance in half the time. Philoponus reported that their falling times differed only slightly, and that with one object twice the weight of the other, the difference could be imperceptible. He did not conclude that weight makes no difference at all; he still held that heavier bodies generally fell faster. But he identified a prediction that observation contradicted, without yet having a complete replacement for the theory. Philoponus’ criticism survived because other people copied, studied, and translated his works. Galileo knew and discussed his arguments, although the extent of their direct influence on his discoveries remains uncertain. The scholarly traditions that preserved Aristotle’s work also preserved objections to his ideas. Both carefully arranged tests and the sharing of criticism therefore existed long before the seventeenth century. What required further development was the organization of these activities into a sustained practice. Preserving an account allows someone else to examine and repeat the investigation, but that opportunity matters only if people actually take it up. Regular meetings, correspondence and eventually scientific journals helped investigators find one another’s work and gave them reasons to test it.

The Royal Society offers an example of how those arrangements developed. Founded in London in 1660, it grew out of informal gatherings of people interested in investigating nature and applying what they learned to practical problems. Members met regularly to conduct demonstrations and discuss results, while correspondence brought observations from farther afield. Its motto, Nullius in verba, is usually translated as “take nobody’s word for it.” Regular meetings gave that ambition a practical form: an investigator could describe an experiment before people who could question both the procedure and the interpretation. Henry Oldenburg, the Society’s secretary, maintained an extensive network of correspondents and in 1665, he began publishing Philosophical Transactions, bringing reports of ongoing investigations to readers who had neither attended the meetings nor received the original letters. This was one of the first scientific journals and represents a major step along the path to the modern scientific process. Making findings publicly available meant that other scientists could consider the ideas, criticize them, and run their own independent tests.

Publication also offered a practical incentive to share, as a dated account could establish who had first reported a discovery. In seeking public credit, an investigator exposed the claim to examination by others, including competitors. Readers could question the methods, attempt to reproduce the result, or propose a different explanation. These arrangements encouraged scrutiny, but publication alone could not guarantee that anyone would check a result or correct an error. Then, as now, our confidence in a finding should grow as it survives examination by people using different methods and approaching the question with different expectations. For someone encountering a claim in a news story or social media video, much of that examination will necessarily have been done by others. My student cannot inspect a data center or repeat every calculation behind its reported water use. But she can learn to ask what was measured, whether others can inspect the methods, whether independent investigators obtain similar results, whether the comparisons are fair, and what the accumulated evidence supports. These five questions help reveal how thoroughly a claim has been checked and how much confidence it has earned and we can apply each of them to her bottle of water. 

Before doing so, we need to determine what type of claim we are examining. Public controversies often combine questions about what is happening, predictions about what will happen, and judgments about what should be done. These questions are connected, but they require different kinds of answers. Climate change, for example, involves both empirical questions and judgments about how to respond. Whether the planet is warming is an empirical question: we can compare temperature measurements taken over time, while checking whether changes in instruments, locations, or methods could account for the apparent trend. What causes the warming is also an empirical question, although answering it requires more than establishing that temperatures have risen. We need to compare possible explanations and ask which best accounts for the observations. Predicting future warming adds assumptions about future emissions and how the climate will respond. Evidence can constrain those predictions, and their accuracy can be assessed as observations accumulate. Deciding what risks to accept, how quickly to act, and who should bear the costs also requires judgments about priorities and obligations. People can agree about the physical evidence and still disagree about those choices. Good decisions require us to understand both the evidence and its uncertainties. My student’s question about AI and water use contains a similar mixture. How much water a particular system uses to complete a specified task is an empirical question. How much water data centers will use in 2030 is a forecast that depends on assumptions about construction, demand, and technological change. Whether that use is excessive requires us to consider what the water would otherwise support and how we value the competing uses. Determining who receives the benefits and who bears the costs involves empirical investigation; deciding whether that distribution is fair requires a further judgment.

Evidence about water use can inform these decisions, but a measurement alone cannot settle them. A lower estimate of water use might weaken one objection to a proposed data center while leaving concerns about the distribution of benefits unanswered. Equally, a company’s poor treatment of a community would not establish that a particular estimate of its water use is accurate. We need to examine each claim closely enough to know what evidence would support it and what remains to be decided.

Question 1: What exactly was measured? My student’s claim sounds specific: one question uses one bottle of water. But “a question” is not a standard amount of computational work, and “water use” can, somewhat non-intuitively, include several different things. Before checking the number, we need to establish which system performed what task, under what conditions, and what the calculation counted. One prominent source of the bottle comparison was a study released in 2023 by researchers at the University of California, Riverside, and the University of Texas at Arlington. Riverside’s announcement described an estimated half-liter of water for approximately twenty to fifty questions and answers. That estimate included water consumed both in cooling the data center and in generating its electricity. The researchers were calculating a footprint from energy demand and water-use information, rather than measuring water flowing through a pipe each time someone entered a question. Even on its own terms, this was a bottle spread across multiple exchanges. A 2024 Washington Post calculation put a 100-word GPT-4 email at 519 milliliters, assuming 140 watt-hours of electricity. Neither calculation establishes a fixed water cost for asking any chatbot anything. The electricity required is one input; the water consumed per unit of electricity and the cooling arrangements are others. Changing those inputs changes the answer. To assess an estimate, we need to inspect the assumptions as well as the arithmetic.

Google reported that the median Gemini Apps text prompt in May 2025 consumed an estimated 0.26 milliliters of water, or about five drops. It derived this estimate from measurements of its serving infrastructure’s energy use and the previous year’s average of data-center water consumption per unit of energy. This water figure covered onsite consumption, excluding water consumed in generating electricity. It described the median text prompt, rather than the average across all requests or the demands of image and video generation. That result provides evidence that onsite water consumption for a text response can be very small. However, it does not establish the full water footprint of every chatbot request. One possibility is that some of the difference between Google’s figure and the earlier estimates are the result of efficiency gains. However, we can’t easily make this comparison because the systems, tasks, dates, and accounting methods differ. These differences also do not make every estimate equally credible. We still need to examine how well the inputs are supported. The answer depends substantially on what the number describes. Are we counting only onsite water consumption, or electricity generation too? One specified task, a typical request, or an entire conversation? Once the claim is clearly defined, we can ask whether its methods and supporting evidence are available for anyone else to examine.

Question 2: Can other people inspect it? Once we are clear on what a claim means, we need to examine how its authors arrived at it. What did they observe directly? What did they calculate or assume? Could someone else follow the procedure and identify a mistake? The public record gives people something more substantial to examine than the author’s assurance that the work was done carefully. Google’s technical paper illustrates both the usefulness and the limits of disclosure. It describes how the researchers combined energy measurements with water-use information, allowing readers to identify what the estimate includes and what it leaves out. But it does not provide outsiders with the underlying operational records needed to reproduce the full calculation independently. We can examine the reasoning while recognizing that some of the evidence remains under the company’s control. For a company reporting on its own environmental impact, that distinction matters: it has access to valuable information and an obvious financial interest in how the findings are understood.

A reader of a published explanation might find an arithmetic error, question an assumption, or notice that a conclusion extends beyond the systems actually studied. Access to the underlying data permits further questions about how observations were selected and whether the reported summary represents them fairly. Disclosure creates opportunities for these different kinds of checking. In 2021, an Oregonian reporter requested records of Google’s water use in The Dalles, Oregon. The city withheld them on trade-secret grounds and, after the district attorney ordered their release, sued to prevent disclosure. A settlement in December 2022 required the city to release annual water-use records for 2012–2021 and to provide comparable information in response to future requests. Residents’ ability to examine the company’s demands on their water system depended on a reporter pursuing the records and a legal process securing access to them. Those records could not settle every question about whether a data center benefited the community. They could, however, establish part of the factual basis for that discussion. Without them, residents had fewer ways to assess assurances about resource use or allegations of harm. Withholding evidence does not establish that a claim is false, but it limits the confidence outsiders can reasonably place in it.

My student need not personally audit every source she encounters, but she can ask whether the evidence is available to people capable of examining it. That leaves a further distinction to consider: making scrutiny possible is one achievement; having independent investigators actually undertake it is another.

Question 3: Has anyone independent checked it? A claim can appear in a newspaper, dozens of videos, and hundreds of social media posts while still resting on one calculation that hasn’t been independently verified. Independent checking requires someone to examine the work, repeat the investigation, or approach the question through an alternative method that could reveal errors in the original result. Peer review, while far from perfect, provides an opportunity for some of this investigation. My students often leave the textbook explanation of the scientific method without understanding the difference between an article that has been gone through by an editor and one subject to peer-review – examined by researchers with relevant expertise in a field pertinent to the claims. In a peer-reviewed research journal, reviewers assess whether a paper’s methods and results support its conclusions. They can identify overlooked alternative explanations, questionable analyses, and claims that go beyond the evidence. But reviewers usually examine the work as presented to them; they do not repeat the full investigation.

Important errors may therefore become apparent only when someone attempts the investigation again. In 2015, the Open Science Collaboration attempted to replicate one hundred studies from psychology journals. While 97 percent of the original studies reported statistically significant results, shockingly – only 36 percent of the replications did. Furthermore, the replicated effects were only about half as large on average. A result falling short of statistical significance does not by itself establish that no effect exists, but these findings showed how much uncertainty could remain after publication and peer review. 

Independence is important because investigators can share mistakes as well as insights. With regard to my student’s question, several calculations that all assume the same electricity demand could agree closely even if that assumption were wrong. Confidence grows more when different approaches, with different opportunities for error, produce compatible results. Repeating a procedure can show whether its result recurs, whereas approaching the problem with a different procedure can help establish whether the result depends on an unnoticed feature of the original setup.

Regarding datacenter water use, there are independent and somewhat converging results, though the full picture remains a bit muddy. In February 2025, an analysis published by Epoch AI estimated roughly 0.3 watt-hours for a typical GPT-4o text query, using assumptions about the model’s computational requirements and the hardware serving it. Google subsequently reported 0.24 watt-hours for its median Gemini Apps text prompt, based on measurements within its own infrastructure. These concerned different systems and were not replications of the same experiment. Nevertheless, their rough agreement nevertheless provides some support for the narrower conclusion that ordinary text requests can require only a few tenths of a watt-hour. We can now return to the electricity assumption behind the Washington Post’s bottle-per-email estimate: 0.14 kilowatt-hours, or 140 watt-hours. That is roughly five hundred times the electricity in these later estimates. The systems, tasks and dates differ, so this comparison does not necessarily establish that the original calculation was wrong for the task it described. It does show why we should not carry that assumption forward as the electricity required for any chatbot response. However, this is still only evidence about electricity. To calculate the associated water consumption, we need to know how a particular data center is cooled and how its electricity is generated.

For my student who is frequently exposed to claims of very high water usage by AI, the practical question is whether the sources repeating a number have contributed any additional checking. Did someone inspect the assumptions? Obtain new measurements? Try another method? Ten articles tracing back to the same study still leave us with one study. What increases confidence is additional work that had a meaningful chance of finding something wrong.

Question 4: Does the comparison support the conclusion? Earlier, we saw how ordinary observation can appear to confirm incorrect explanations. Numbers can mislead us in a similar way, even when they are accurate. Hearing that a data center consumes millions of liters of water sounds alarming, especially when I compare it with my personal water use. But most of us have little intuitive understanding of water use at an industrial scale, or how AI datacenter water use compares with other industries. We need a comparison appropriate to the question before we can decide what the number means. A per-prompt estimate tells us about the resources required for a task, whereas an annual total tells us about aggregate demand. If each request uses half as much water but the number of requests quadruples, total consumption doubles. Improving efficiency and increasing consumption can therefore occur together.

We also need to use consistent definitions to enable accurate comparisons. Water withdrawal means taking water from a source; some of it may subsequently be returned. Water consumption is the portion unavailable for immediate local reuse, often because it evaporates. Comparing one industry’s consumption with another’s withdrawals can create a large apparent difference simply by changing what is counted. 

Lawrence Berkeley National Laboratory estimated that United States data centers directly consumed about 66 billion liters of water in 2023, with nearly 800 billion additional liters attributed to generating their electricity. These modeled estimates cover all data centers, including uses beyond AI, and imply a combined operational footprint of roughly 870 billion liters that year. For comparison, the U.S. Geological Survey estimated that irrigation in the United States consumed approximately 100 trillion liters of water in 2015. That is more than one hundred times the estimated annual data-center footprint, even including electricity generation. The estimates concern different years and methods, so this comparison establishes their approximate relative scale rather than a precise current ratio. This national comparison still cannot answer the questions faced by a particular community. Water available elsewhere in the country does not relieve a shortage in the town’s own watershed, and a facility’s small national share does not establish that its local effects are small. For my student, the question is whether the comparison answers her concern. A few drops per request cannot settle the question of an industry’s total demand. A large annual total cannot establish the consequences for a particular community. We need to choose the comparison according to what we want to find out, and keep that question fixed when interpreting the result.

Question 5: What survives when evidence accumulates? As evidence accumulates, we can begin to ask more demanding questions of an explanation. Does it account for findings from different populations and methods? Does it explain when an effect appears, how its strength varies, and when it disappears?

The evidence linking smoking to lung cancer illustrates this process. Richard Doll and Austin Bradford Hill published an influential study in 1950 comparing the smoking histories of patients with and without lung cancer. They subsequently began following British doctors over time, reporting initial mortality findings in 1954. The different designs approached the question from opposite directions: one looked back from illness to previous exposure, while the other recorded smoking habits before observing subsequent deaths.

The statistician Ronald Fisher questioned whether smoking itself caused the cancer. Perhaps an inherited predisposition made people both more likely to smoke and more susceptible to the disease. Fisher also worked as a consultant to the tobacco industry, which promoted this alternative explanation. His financial relationship deserved scrutiny, but the possibility of a common cause was a question that evidence needed to address. The causal link between smoking and cancer gained strength as findings accumulated – the association appeared across many studies; risk increased with the amount and duration of smoking; and people who stopped smoking subsequently had lower risks than those who continued. Any competing explanation had to account for these patterns and Fisher’s explanation did not. The accumulated evidence undermined the claim that an inherited predisposition could explain the association without smoking itself causing harm. Genetic differences could influence susceptibility without explaining away the harmful effects of smoking.

Returning to my student’s question, the evidence supports several conclusions with different degrees of confidence. The claim that every chatbot question consumes a bottle of water cannot be sustained as a general rule. Google’s study provides evidence that onsite water consumption associated with a text prompt can be a fraction of a milliliter, although that estimate excludes water consumed in generating electricity and does not describe every kind of AI task. Could the omitted electricity-related water bring the figure back up to a bottle? We can explore that possibility with a calculation, provided we make its assumptions clear. Google reported 0.24 watt-hours for its median text prompt. If we apply Berkeley Lab’s estimated U.S. data-center electricity factor of 4.52 liters of water per kilowatt-hour, generating that electricity would consume about 1.1 milliliters. Adding Google’s onsite estimate gives about 1.3 milliliters in total. This is an illustrative calculation, not a measurement of Gemini’s full footprint: the electricity supplying Google’s systems may differ from the national average used here. Under this assumption, however, including electricity generation increases the estimate several times over while leaving it far below a 500-milliliter bottle. The omission matters, but it does not by itself rescue the bottle-per-question claim. We still cannot assign a precise footprint to my student’s particular request.

At the industrial scale, the Berkeley Lab estimates indicate growth in direct water consumption between 2014 and 2023. Our comparison with irrigation places data-center demand in a broader national context. Neither finding establishes what a particular community can sustainably supply. That requires local evidence about water availability, competing demands, and the facility itself. My student can therefore become less confident in the bottle-per-question claim without concluding that every concern about data centers is unfounded. She can also recognize uncertainty about a particular request without treating all estimates as equally plausible. The aim is to reach the most specific conclusions the evidence supports, and to know which questions remain open.

The scientific process described in this essay is extremely powerful and underlies much of human progress – but correction depends on people having the opportunity and incentive to examine claims. It does not happen automatically, and the process can be lengthy and complicated. My student is not training to become a scientist. I cannot expect her or the average citizen to inspect every claim they encounter, look through detailed methods and primary sources, and seek out additional evidence from peer-reviewed research. Even understanding one paper can require substantial expertise, and knowing how it fits with everything else published on the subject requires a great deal more. Scientists themselves depend on other researchers to gather and evaluate that larger record. So we’re left with a practical question. When my students are confronted with issues like these, where should they actually look for more information?

Different kinds of publications serve different purposes. Understanding these can help us to decide where to focus when looking for information. A primary research paper reports a particular investigation. A preprint makes a manuscript available before formal peer review. Neither, by itself, tells us how well the findings have survived subsequent scrutiny. These are not good places to start for someone outside of the field who is trying to get a good overview or understanding of the consensus positions. A systematic review addresses a defined question by searching for relevant studies according to explicit rules, assessing their limitations, and considering what they establish together. A meta-analysis can combine sufficiently comparable results statistically. These methods help reviewers examine a body of evidence without simply selecting the studies they find most persuasive. Their reliability still depends on the search, the selection rules, and the quality of the underlying research. These caveats aside, for many questions, an accessible summary of a well-conducted review is therefore a useful place to begin. Cochrane provides plain-language summaries of its health research reviews. National Academies consensus reports assemble expert assessments of scientific and technical questions and undergo independent review. The IPCC’s assessments draw together climate research and explicitly distinguish findings by their degree of confidence. These publications offer readers a way to benefit from work that would be difficult to repeat individually.

Relying on expertise is reasonable for busy humans with finite time and energy. However, we should not blindly trust any organizations based on their perceived authority or reputation alone. What matters is why that expertise deserves our trust: whether conclusions are supported by evidence, exposed to independent criticism, and revised when better evidence arrives. Reputations and practices can change, so the same five questions apply to reviews and reports: what did they examine, what can others inspect, how independently have their conclusions been checked, do their comparisons support what they say, and what does the accumulated evidence establish?

I want my students to leave a science course with more than the ability to select a correct answer on multiple choice exams. I want them to have some understanding of why an answer deserves their confidence, and what evidence might change it.