The carbon cost of one prompt, measured
Every AI query leaves a carbon trail spanning electricity grids, cooling systems and the hardware it runs on. Only two companies have ever disclosed the actual number, and the methodology gap between them exposes why transparency, not just efficiency, is the real policy problem.
When you send a text prompt to a large language model, carbon enters the picture at three distinct points. The first is the electricity consumed by the GPU doing the computation, Scope 1 and 2 emissions depending on whether the data centre generates its own power or buys from the grid. The second is the cooling infrastructure keeping those chips within operating temperature. The third, and most commonly omitted, is the carbon embedded in manufacturing the hardware itself, the GPUs, servers, networking equipment and building materials, Scope 3 emissions that can account for 20 to 30 per cent of a hyperscale facility’s lifetime footprint[1]Guidi et al. (Harvard/Meta)View source →.
Most public estimates capture only Scope 1 and 2 operational emissions, and almost none are based on actual production measurements. The Stanford Foundation Model Transparency Index, published in December 2025, found that ten of the thirteen AI companies assessed disclosed none of the three key environmental metrics[2]Stanford CRFMView source →, energy, carbon or water, with an average transparency score of 40 out of 100. The opacity has structural consequences: regulators cannot set evidence-based limits without baselines, and enterprises with Scope 3 reporting obligations under the EU’s CSRD[3]European CommissionView source → cannot account for their AI usage without supplier-side disclosure. Two companies have stepped out of the silence.
Carbon per prompt: disclosed vs estimated
What the two disclosures actually show
In August 2025, Google published the first production-scale measurement of Gemini inference. The median Gemini Apps text prompt consumed 0.24 watt-hours of electricity and emitted 0.03 grams of CO₂ equivalent[4]Google Cloud BlogView source →, derived from real fleet data and using the company’s fleet-wide average grid intensity. Between May 2024 and May 2025, energy per prompt fell by a factor of 33 and carbon per prompt by a factor of 44, attributed to mixture-of-experts architectures, TPU improvements and 24/7 carbon-free energy procurement, working backwards from a 2024 baseline of roughly 8 watt-hours.
In July 2025, Mistral published the first peer-reviewed full lifecycle assessment of a production model. Conducted with Carbone 4 and independently audited by ADEME across training, hardware manufacture and 18 months of inference, a 400-token response via Le Chat produced 1.14 grams of CO₂ equivalent and 45 millilitres of water[5]Mistral AI / Carbone 4 / ADEMEView source →. Mistral’s figure is roughly 38 times higher than Gemini’s, but these are not measuring the same thing: Google’s excludes embodied hardware and reflects aggressive clean-energy procurement, Mistral’s includes hardware and reflects a more typical grid. Together they bracket a range; neither is sufficient on its own.
The gap between Gemini and Mistral Large 2 is not a measurement error. It is the industry’s transparency problem made visible.
Anatomy of a prompt’s carbon
A per-prompt carbon figure is the product of at least four multiplied variables. Each layer matters because different actors control different levers, from the model developer who picks the architecture to the operator who chooses the hardware refresh cycle.
| Layer | Metric | Typical range | Who controls it |
|---|---|---|---|
| GPU compute | Energy (Wh/prompt) | 0.03 – 1.7+ | Model developer (architecture, size) |
| Facility overhead | PUE multiplier | 1.06 – 1.58 | Data-centre operator (cooling design) |
| Grid carbon | Intensity (g CO₂e/kWh) | 20 – 650 | Cloud provider (energy procurement) |
| Embodied hardware | Share of lifecycle | 20 – 30% | Operator (hardware refresh cycles) |
| Net Scope 1+2+3 | g CO₂e/prompt | 0.03 – 4.32+ | All of the above, simultaneously |
The grid variable, a 10×+ swing from a single choice
Of the four layers, grid carbon intensity produces the widest variation. A Harvard study of 2,132 US data centres found their electricity-weighted average carbon intensity was 548 g CO₂e per kWh, about 48 per cent higher than the US national average[9]Guidi et al. (Harvard T.H. Chan)View source →, because data centres cluster in carbon-intensive regions; Virginia’s PJM, which hosts the most US data centres, runs at roughly 576 g. The contrast with cleaner grids is stark: France’s nuclear grid sits near 55 g and Nordic hydro and wind near 20[10]EPA eGRID 2023View source →, more than a 28-fold difference for identical compute. Grid location is the single largest lever available without changing a line of model code.
Grid carbon intensity by region
Embodied hardware, the hidden 20–30 per cent
When an operator buys a rack of H100 GPUs, the emissions from chip fabrication, rare-metal extraction and component manufacturing are locked in before the hardware runs a single query. Embodied emissions account for 20 to 30 per cent of a hyperscale facility’s lifetime footprint, and that share is growing[1]Guidi et al. (Harvard/Meta)View source →as operational grids get cleaner, the manufacturing phase becomes the dominant remaining source. Google’s own hardware LCA found compute carbon intensity improved three-fold from TPU v4i to TPU v6e[11]Google / arXivView source →, with cascaded reuse offsetting embodied emissions by up to 4 per cent. For operators on commodity hardware with shorter refresh cycles, the Scope 3 burden is proportionally larger.
Training versus inference: where lifetime carbon accumulates
The public conversation has focused disproportionately on training. GPT-3’s training run, 1,287 MWh and about 552 tonnes of CO₂ equivalent[12]Patterson et al. (2021)View source →, became the reference figure in almost every policy discussion. That no longer reflects where emissions go. The breakeven is mechanical: at ChatGPT’s reported 2.5 billion prompts a day and 0.42 watt-hours per GPT-4o query, cumulative inference energy matches GPT-4’s estimated 50,000 MWh training energy in about 48 days[6]Hao et al.View source →. Across a model’s lifetime, inference can account for over 90 per cent of total operational energy[14]Luccioni et al. (2024)View source →.
| Metric | Figure | Detail | Source |
|---|---|---|---|
| Training energy (GPT-4, est.) | ~50,000 MWh | ≈25,000 A100s over 90–100 days | [13] |
| Inference breakeven | ~48 days | 2.5B prompts/day × 0.42 Wh per query | [6] |
| Inference share of operational energy | >90% | at scale, over the model’s lifetime | [14] |
| Global data-centre demand, 2030 | 945 TWh | IEA, roughly Japan’s total electricity use | [15] |
The disclosure gap is the story
The per-prompt figures that exist, 0.03 g for Gemini, 1.14 g for Mistral Large 2, 0.15 g for GPT-4o by third-party estimate, are not directly comparable because they measure different things. None covers full Scope 3 supply-chain emissions, and no methodology is settled at industry level. Yet the demand is regulatory, not optional: even after the February 2026 Omnibus simplification narrowed its scope, the CSRD still requires Scope 3 disclosure from Wave 1 companies now and Wave 2 from 2028[3]European CommissionView source →, and AI usage sits in Scope 3 for every enterprise that does not own its compute.
There are deployable levers already. Carbon-aware workload scheduling, routing non-urgent inference to greener regions, can cut emissions by 14 to 41 per cent[17]AICompetence / USENIX NSDI 2025View source →with no change to the model itself. And Mistral’s ISO 14040/44-audited LCA shows full measurement is possible at production scale with peer review. But Google’s own emissions still rose 51 per cent between 2019 and 2025[16]CNaughtView source → as volume outran efficiency, and the IEA projects global data-centre demand reaching 945 TWh by 2030[15]IEAView source →. The question is not whether the industry can measure these emissions, Mistral proved it can. It is whether the rest of the industry is required to follow.
References
- Guidi et al. (Harvard/Meta), Embodied emissions: 20–30% of hyperscale lifetime footprint (2024)https://sustainableatlas.org/post/data-story-data-center-energy-consumption-water-carbon-trends-1624
- Stanford CRFM, Foundation Model Transparency Index (December 2025)https://crfm.stanford.edu/fmti/
- European Commission, CSRD Omnibus I simplification package, Directive (EU) 2026/470https://finance.ec.europa.eu/financial-markets/company-reporting-and-auditing/company-reporting/corporate-sustainability-reporting_en
- Google Cloud Blog, Measuring the environmental impact of AI inference (August 2025)https://cloud.google.com/blog/products/infrastructure/measuring-the-environmental-impact-of-ai-inference
- Mistral AI / Carbone 4 / ADEME, Lifecycle Assessment of Mistral Large 2 (July 2025)https://mistral.ai/news/our-contribution-to-a-global-environmental-standard-for-ai
- Hao et al., Energy Efficiency Benchmarks for LLM Inference (2024)https://arxiv.org/abs/2505.09598
- Smartly.AI, CO₂ per ChatGPT query: 2.5–5 g range; 4.32 g detailed estimatehttps://smartly.ai/blog/the-carbon-footprint-of-chatgpt-how-much-co2-does-a-query-generate
- Jegham et al. (arXiv, May 2025), How Hungry is AI? Benchmarking Energy, Water, and Carbon of LLM Inferencehttps://arxiv.org/abs/2505.09598
- Guidi et al. (Harvard T.H. Chan), Environmental Burden of US Data Centers (US avg 548, Virginia PJM 576 g CO₂e/kWh)https://arxiv.org/abs/2411.09786
- EPA eGRID 2023, US national average grid carbon intensity: 352 g CO₂e/kWhhttps://www.epa.gov/egrid
- Google / arXiv, Life-Cycle Emissions of AI Hardware: A Cradle-To-Grave Approach (Feb 2025)https://arxiv.org/html/2502.01671v1
- Patterson et al. (2021), Carbon Emissions and Large Neural Network Traininghttps://arxiv.org/abs/2104.10350
- Ludvigsen, The carbon footprint of GPT-4 (training energy estimate >50,000 MWh)https://towardsdatascience.com/the-carbon-footprint-of-gpt-4-d6c676eb21ae/
- Luccioni et al. (2024), Inference constitutes over 90% of total operational energy at scalehttps://arxiv.org/abs/2311.16863
- IEA, Energy and AI report (2025): global data centre demand 945 TWh by 2030https://www.iea.org/reports/energy-and-ai
- CNaught, How Much Carbon Does AI Actually Use? Google +51%, Microsoft +29% (Jan 2026)https://www.cnaught.com/blog/how-much-carbon-does-ai-actually-use-and-why-its-so-hard-to-find-out
- AICompetence / USENIX NSDI 2025, Carbon-aware scheduling: 14–41% reductions via temporal and spatial shiftinghttps://aicompetence.org/cut-ais-carbon-footprint-smart-green-scheduling/