Almost everyone answering that question is working from a feeling — that AI is ruinous for the planet, or that it is a rounding error. Both camps cite real numbers; the numbers are between one and two orders of magnitude apart. mint trees is an attempt to get past the feeling: measure one real AI product end to end, in the open, with the method and the error bars on the table.
No product. No credits for sale.
A measurement problem, worked in public.
bdrop.studio is the first case, not the point.
Method and factors reusable under CC BY.
Two teams can measure the same web app and land a factor of ten apart — not because one lied, but because one counted only the server and the other counted the hardware, the network, the idle capacity and the laptops. Here is the line we draw, and what sits on each side of it.
What the software actually draws, including the data centre's overhead. Measurable for our own machines, modelled for managed services.
How dirty that electricity was — which grid, which hour. The only term you can change without touching a line of code.
The hardware's manufacturing footprint, amortised across its useful life and the share of it you occupied.
Per what? Per request, per user-month, per deploy, per shipped feature. Change R and the headline number changes with it.
Pick a monthly electricity figure for a workload and watch the same kilowatt-hours land fifteen times apart, purely by region. Annual average grid intensities, 2024. 4
Equivalences convert a figure you have measured into one you have not. A “tree-year” depends on species, age, climate and what you assume about the tree’s next fifty years; a car-kilometre depends on a fleet average that differs by a factor of two between test cycle and road. Both add uncertainty while appearing to add clarity. If a number needs an analogy to be persuasive, the honest move is to explain the number.
bdrop.studio runs on managed infrastructure, like most small teams. That means most of our footprint sits inside someone else’s reporting boundary. Here is the honest state of each line.
This is an instrumentation register, not a result. We publish no figures for our own stack yet — only what each line would take to measure and how far we are from being able to. Numbers appear here when they come from telemetry we can show.
Inference during development: long contexts, tool calls, repeated runs. Per-prompt medians published by providers are small, but agentic sessions are orders of magnitude above a median chat prompt.
Missing: per-session token and energy telemetry from providers. We estimate from token counts and published per-prompt figures.
Build minutes are billable, therefore countable. Runner class, region and duration give a defensible energy estimate — and preview deploys on every push add up faster than production.
Missing: runner hardware spec and regional placement. Embodied share of the runner fleet is modelled, not measured.
Serverless and edge invocations, bandwidth, image optimisation. Duration × memory class is a usable proxy for energy once a regional grid factor is applied.
Missing: idle and redundancy overhead, and a location-based figure — provider reporting is market-based.
The one provider giving us region-level carbon data per service, including a location-based view. Our best-instrumented line by a wide margin.
Missing: embodied hardware attribution per workload, and hourly rather than annual grid resolution.
A database is the quiet constant: it draws power at 3 a.m. with zero users. Instance size × hours is the honest baseline, not query counts.
Missing: underlying host utilisation and multi-tenant allocation. Storage at rest is estimated from disk size.
Laptops, displays, phones — dominated by manufacturing, not use. Amortised over a realistic replacement cycle, this rivals everything above it for a small team.
Missing: verified per-device LCA data beyond manufacturer product sheets.
Use them on us too. If an answer here is missing, that is a finding, and it belongs in the log.
Per request, per user, per month, per shipped feature — a number without a denominator cannot be compared to anything. Watch for totals quoted where a rate belongs, and rates quoted where a total belongs.
Market-based accounting subtracts renewable certificates and power purchase agreements, so a data centre on a coal-heavy grid can report near zero. Location-based reflects the physical electricity actually consumed. Credible reporting shows both; Google's per-prompt figure, for example, is market-based.
Manufacturing is a large, front-loaded share of hardware’s lifetime emissions — for end-user devices it commonly dominates the use phase. If a study only counts electricity, it is measuring part of the problem and calling it the whole.
Attributional asks what share of existing emissions belongs to you. Consequential asks what changes in the world if you stop. They answer different questions and routinely disagree — a marginal kilowatt-hour is usually dirtier than the grid average.
Digital footprint estimates commonly carry a factor of two or more of uncertainty, driven by utilisation assumptions, PUE, allocation and data vintage. A single decimal-precise figure with no range is a presentation choice, not a measurement.
IEA, Energy and AI (2025): data centres ≈ 415 TWh in 2024, ~1.5 % of global electricity, +12 %/yr, projected ~945 TWh by 2030.
iea.org/reports/energy-and-aiGoogle (2025), Measuring the environmental impact of delivering AI at Google scale: median text prompt 0.24 Wh, 0.03 gCO₂e, 0.26 mL water — market-based, excludes training and network.
arxiv.org/abs/2508.15734Green Software Foundation, Software Carbon Intensity specification, standardised as ISO/IEC 21031:2024.
sci.greensoftware.foundationGrid intensities: Ember Electricity Review and national TSO data, annual averages 2024. Indicative — hourly values swing by a factor of three within a single day.
ember-energy.orgWu et al. (2026), Energy use of AI inference, efficiency pathways, and test-time scaling, Joule: frontier models on H100 nodes, median 0.31 Wh/query (IQR 0.16–0.60) — widely cited estimates overstated by 4–20×.
cell.com/jouleEpoch AI (2025), How much energy does ChatGPT use? — model-based estimate of ~0.3 Wh for a typical prompt, with the calculation shown.
epoch.aiFigures on this page are order-of-magnitude estimates for public education, not a verified inventory. Where we state a number from our own stack it is labelled with its confidence. Corrections are welcome and will be published with attribution.
The public argument about AI and the climate is running far ahead of the evidence, in both directions. What is missing is not more opinion — it is a small number of workloads measured end to end by people willing to publish their assumptions and be wrong in public. We can supply exactly one of those. It happens to be ours.
A small generative-AI product — text, image and audio models, managed infrastructure, a tiny team. Unremarkable, which is the useful part: most AI products look like this, and almost none of them are measured.
The boundary definition, the factor registry with its tiers and sources, the uncertainty propagation, the error budget that names which unknown is currently costing you the most. Those are portable. Our kilowatt-hours are not, and nobody should reuse them.
A measured figure that lands outside our published band, or a provider disclosure that collapses one of our wide terms. Either would be a good day. Revisions get published with the reasoning — starting with our own first guess, which was more than ten times too high and is on the record as such.
This log takes no position on whether AI is good or bad for the climate. It is trying to establish what one unit of it costs, to a precision that would let someone else have that argument with evidence instead of adjectives.
Two weeks of a real bdrop.studio project — every agent call, CI minute, preview deploy and idle database hour, measured and published with the raw data alongside.