
In 2023, a widely quoted research paper gave conversational AI a vivid water bill.
A 500 mL bottle, the researchers estimated, could account for roughly 10 to 50 medium-length responses from GPT-3, depending on where and when the model ran.
Two years later, a Google-authored preprint produced a strikingly different figure. The median text prompt sent to Gemini Apps in May 2025 consumed an estimated 0.26 mL of water, roughly five drops.
A bottle for a few dozen responses. Five drops for one prompt.
Those figures sit orders of magnitude apart. They have appeared in public discussion as though one must disprove the other. They do not measure the same thing.
The GPT-3 estimate was modeled from public information and assumptions about energy use, data-center cooling, electricity generation, location, and time. It included water consumed directly at the data center and water associated with generating its electricity.
The Gemini estimate came from Google’s internal production telemetry. It covered a different generation of models, hardware, software, and serving infrastructure, then applied Google’s fleet-level Water Usage Effectiveness to estimate freshwater consumed for data-center cooling. It excluded the water footprint of electricity generation and hardware manufacturing. This is valuable first-party evidence, but the underlying operational data are controlled by Google and cannot be independently reproduced in full.
The gap is not simply a dispute over arithmetic. It exposes a deeper problem with one of the most popular questions about AI infrastructure:
How much water does one AI prompt use?
A prompt is something a person sees on a screen. It is not a stable unit of computing.
Asking a language model to correct one sentence and asking it to analyze a long legal document are both prompts. They do not require the same amount of computation.
A model may process a few words or tens of thousands of tokens. It may produce one sentence or several pages. It may answer directly, search uploaded material, call external tools, generate an image, or spend additional computation on multi-step reasoning.
The model behind the interface matters too. A small model built for routine classification does not behave like a large general-purpose model. A model that activates only part of its architecture for each request may use hardware differently from one that activates the whole network. Quantization, caching, speculative decoding, and other techniques can reduce the work required to produce a response.
Then there is the production system. Providers do not usually dedicate one accelerator to one person’s prompt. Requests can be grouped into batches so expensive hardware serves several users at once. Higher utilization spreads the energy consumed by the machine across more useful work. Low utilization, strict latency targets, and spare capacity kept ready for sudden demand push energy per prompt upward.
A 2025 benchmarking preprint covering 30 language models estimated more than a 70-fold difference in energy consumption between some models for long prompts. Changing batch size also altered energy per request substantially. The work relied partly on public API behavior, hardware specifications, and inferred deployment assumptions rather than direct provider telemetry, so its model rankings should not be treated as universal measurements. It still shows how much model choice, response length, and serving configuration can change the answer.
Google’s own preprint reveals the same problem from inside a production system. A narrow calculation based largely on active accelerator energy produced an estimate of 0.10 Wh for a median Gemini prompt. Once the researchers included host processors and memory, machines kept idle for availability, and data-center overhead, the figure rose to 0.24 Wh.
The model and prompt did not change. The measurement boundary did.
A more complete boundary made the same prompt appear 2.4 times more energy-intensive. Because the water estimate was partly derived from energy use, the water result changed with it.
A per-prompt number should therefore arrive with a description of the prompt, model, hardware, utilization, and method. Without them, the unit looks far more concrete than it is.
AI infrastructure can be associated with water at several points, and a study may include one, two, or all of them.
The most visible is water used directly at the data center. Computing equipment turns electricity into heat. The cooling system has to move that heat away from the processors and eventually out of the facility. Some systems reject heat through evaporation, consuming water that must be replaced. Others circulate liquid through closed circuits or reject heat through outdoor air with little continuing freshwater input at the site.
The complete cooling path runs from the processor to final heat rejection, while closed-loop cooling requires a separate distinction between circulation and consumption.
AI makes cooling harder because it can concentrate more electrical power into each rack. ASHRAE recommends liquid or liquid-assisted cooling for high-density AI and high-performance computing systems while retaining air cooling for equipment that still needs it. Higher density changes how heat is captured and transported, but it does not dictate how the facility finally rejects that heat or how much freshwater it consumes.
Water also sits upstream. Electricity generation can withdraw and consume water through power-plant cooling and other processes. The amount depends on the technology, region, and electricity mix serving the facility. An AI system with little direct cooling-water consumption may therefore retain an indirect footprint through its power supply.
The GPT-3 estimate included both direct cooling water and electricity-related water. Google’s five-drop estimate used data-center WUE to estimate direct cooling-water consumption but did not add the water used to generate electricity. That difference alone makes the two figures unsuitable for direct comparison.
A third layer appears before servers reach the data center. Semiconductor fabrication requires ultrapure water, while manufacturing servers, cooling equipment, electrical systems, and buildings creates additional water and material demands. Public information about this embodied water remains much thinner than operational data. Lifecycle assessment widens the boundary beyond daily operation.
When an estimate says an AI prompt “uses water,” the first question should be which water it includes: direct cooling, electricity generation, hardware manufacturing, construction, withdrawal, or consumption. The broader data-center water question depends on keeping those categories apart.
Early discussion of AI’s environmental footprint concentrated heavily on training. That made sense. A major training run brings thousands of accelerators together for an intensive, measurable period and produces the kind of large number that can be attached to one model release.
But the final run is only one part of creating and operating an AI system.
Researchers test architectures, tune parameters, prepare data, run failed experiments, evaluate checkpoints, and repeat training at different scales before the final version exists. Much of this development work is rarely included in public model disclosures.
A 2025 preprint presented as an ICLR spotlight measured the development and training of a series of language models ranging from 20 million to 13 billion active parameters. Across that specific program, the researchers estimated 2.769 million liters of water when hardware manufacturing, model development, and final training were included. Model development accounted for roughly half the impact of the final training runs.
Those findings describe one program, not a universal ratio for AI. Their significance lies in showing how much work can disappear when reporting begins with the final run.
After deployment comes inference: the repeated work of generating responses, classifications, recommendations, images, or predictions for users.
Training may happen in a finite series of campaigns. Inference happens every time the system is used. A lightly used model may never accumulate an inference footprint comparable with its development and training. A service handling enormous demand may eventually spend far more energy serving users than it spent creating the model.
There is no fixed training-versus-inference split. It depends on model lifetime, adoption, request complexity, hardware efficiency, utilization, and how often the system is updated or replaced.
Google’s 0.26 mL figure sounds reassuringly small. For the median Gemini text prompt measured in May 2025, it may be a sound estimate within the boundary Google defined.
But a median does not describe every request.
Google chose the median because prompt-level energy use was highly skewed. A smaller group of long requests and models with lower utilization consumed disproportionately more energy. An arithmetic average would have been pulled upward by those outliers, while the median better represented the typical prompt in the measured product mix.
That makes the figure useful for one question: what did a typical Gemini Apps text prompt look like in that system and period?
It is much weaker as the basis for calculating total service water use. Multiplying the median by the total number of prompts does not produce an accurate fleet total when the distribution is skewed. Nor can the number be transferred automatically to another model, provider, region, cooling system, or year.
This is where tiny units can create false reassurance or exaggerated alarm. A commentator can multiply a large per-prompt estimate by an assumed global volume and produce a frightening annual total. A provider can publish a small median prompt and make aggregate infrastructure appear negligible.
Both calculations may use real numbers. Neither necessarily describes the whole system.
For the community hosting an AI data center, the more useful figures operate at another scale: annual water withdrawal and consumption, peak summer demand, the water source, cooling behavior under full load, and the water footprint of the electricity supply.
A per-prompt estimate can help model developers track efficiency. It cannot replace facility-level reporting. The same principle applies to Water Usage Effectiveness and Power Usage Effectiveness. Ratios become useful when their boundaries, locations, operating conditions, and time periods remain visible.
A serious per-prompt estimate begins by defining the work: the model and version, type of request, input and output length, modality, and whether additional reasoning or tool use occurred. It should say whether the result represents a median, average, benchmark, or measured production request.
The infrastructure description matters just as much. Hardware type, host-system energy, utilization, batching, idle capacity, data-center overhead, and cooling method all affect the result.
Then comes geography. A water estimate should identify the data center or regional mix behind the calculation, the reporting period, cooling-water intensity, and local electricity supply. Annual averages can improve comparison, while seasonal and peak figures may be more useful for understanding local pressure.
Finally, the boundary has to be named plainly. Direct site consumption should not be confused with electricity-related water or embodied water. Withdrawal should not be presented as consumption. A cooling-only result should not be described as the total water footprint of AI.
Berkeley Lab’s 2025 review shows why this detail matters. Across the conditions studied, workload water use varied by more than 10,000 times. Server efficiency ranked first among the determinants, followed by the water intensity of electricity, utilization, cooling, infrastructure efficiency, climate, inactive servers, and replacement cycles.
A trustworthy number does not need to cover every stage perfectly. It does need to say what has been left out.
Policloud develops and deploys physical, modular data-center infrastructure for defined sites. AI training and inference may be possible workloads where the configuration supports them, but the word “AI” does not determine one infrastructure design or one water result.
A model does not arrive with a fixed amount of water attached to every request. Its footprint emerges from the work it performs and the physical system serving it.
Software determines how much computation happens. Hardware determines how efficiently that computation runs. Utilization determines how much productive work shares the equipment. Cooling determines how heat leaves the facility. Electricity supply adds another geographical layer. Manufacturing extends the boundary beyond operations.
The user sees one prompt box. The water footprint belongs to everything behind it.
So how much water does one AI prompt use?
Five drops may be a defensible estimate for one median prompt in one production system under one stated boundary. A fraction of a bottle may be a defensible modeled estimate for another system using a wider boundary.
Neither number belongs to AI as a whole.
Without the model, workload, infrastructure, place, time, and accounting method, a prompt has no fixed water footprint.
It has only a memorable number.