One coding request on a 7B model running on our own GPU took 238 tokens, 5.65 seconds and 0.38 Wh with retries, about $0.00006 of electricity.
A hard coding request to a frontier API cost $0.0003 to $0.012 at list price, because the models spent 141 to 3,080 output tokens on it.
Energy per request ranged from 0.003 Wh for a 2B classifier to 0.62 Wh for a 27B coding answer; Google's published median Gemini prompt is 0.24 Wh.
Loupe by LeanLM (not affiliated with Google's LearnLM educational AI) measures what AI workloads cost to serve. Every number below comes from our own runs or a named published source.
How do you calculate LLM cost per request?
LLM cost per request is input tokens times the input price, plus output tokens times the output price. Both prices are quoted per million tokens, so divide by 1,000,000 before multiplying.
Worked example: 1,000 input tokens and 500 output tokens on Claude Sonnet 5 ($2/$10 per million). Input costs 1,000 × $2 / 1,000,000 = $0.002. Output costs 500 × $10 / 1,000,000 = $0.005. Total: $0.007. The same request on Gemini 3.1 Flash-Lite ($0.25/$1.50 per million) costs 1,000 × $0.25 / 1,000,000 + 500 × $1.50 / 1,000,000 = $0.00025 + $0.00075 = $0.001, about a seventh of the Sonnet 5 price. That's your cost per API call.
Output tokens usually dominate the bill. They're priced at 5 to 6 times the input rate on the models above, and a model that thinks before answering pays for that thinking in output tokens.
Tokens per request aren't fixed by the task either. On the same hard coding tasks, list-price API calls ran from $0.0003 to $0.01172 because the models spent very different amounts of output on the same problem: gpt-6-astra wrote 141 tokens per call, gpt-6-luna wrote 557, and gemini-3.8-flash wrote 3,080. Price per token alone doesn't tell you price per request; you need the token count too.
What one request costs, measured
Measured, one request cost between $0.0000004 of electricity for a 2B classifier and $0.01172 at list price for the most verbose frontier answer:
| Request | Model | Output tokens | Seconds | Joules | Wh | Cost |
|---|---|---|---|---|---|---|
| Coding task, one attempt | Qwen2.5-Coder-7B | 728 | 0.20 | $0.00003 electricity | ||
| Coding task, feedback retries | Qwen2.5-Coder-7B | 238 (mean) | 5.65 (mean) | 1,367 (mean) | 0.38 (mean) | $0.00006 electricity |
| Coding task, blind retries | Qwen2.5-Coder-7B | 182 (mean) | 1,152 | 0.32 | $0.00005 electricity | |
| Coding task, one attempt | Qwen3-Coder 30B (MoE) | 221 (median) | 1.8 (median) | 448 | 0.12 | $0.00002 electricity |
| Coding task, one attempt | Qwen 3.8 27B | 482 (median) | 6.7 (median) | 2,221 | 0.62 | $0.00009 electricity |
| ContractNLI classification | 2B model | 9.0 | 0.0025 | $0.0000004 electricity | ||
| Banking77 classification | 2B model | 24.3 | 0.0067 | $0.0000010 electricity | ||
| Hard coding task | gpt-6-luna (API) | 557 (mean) | 5.9 (mean) | $0.0003 list price | ||
| Hard coding task | gpt-6-astra (API) | 141 (mean) | 4.1 (mean) | $0.00918 list price | ||
| Hard coding task | gemini-3.8-flash (API) | 3,080 (mean) | 10.6 (mean) | $0.01172 list price |
Local rows: GPU-only energy, idle power subtracted, one AMD Radeon AI PRO R9700, one request at a time (batch size 1). Classification rows: model-serving energy, batch size 1, one reference server, published 2026-09-17 on /results. API rows: list price × tokens reported by the provider API, run 2026-09-27; provider-side energy isn't visible to us. Electricity priced at $0.15/kWh; excludes hardware purchase, cooling and the rest of the machine. See the hardware side of the cost. Joules are the mean per task or call; 1 Wh = 3,600 J. Full data: llm-cost-per-request-measured.csv.
How much energy does an AI query use?
A typical published figure for one text query is 0.24 to 0.34 Wh; our own requests ranged from 0.003 Wh to 0.62 Wh depending on model size and task. Energy per ChatGPT query, per Epoch AI's estimate, is about 0.3 Wh.
- Google, August 2025 (Elsworth, Patterson, Dean et al., arXiv 2508.15734; Google Cloud blog): median Gemini Apps text prompt uses 0.24 Wh, 0.03 gCO2e and 0.26 mL water. That's the full serving stack: host CPU and memory, idle provisioned capacity, and data-center overhead (power usage effectiveness, PUE). Counting active machines only gives 0.10 Wh, which Google itself calls the optimistic number.
- Epoch AI, February 2025 (Josh You, "How much energy does ChatGPT use?"): about 0.3 Wh for a typical GPT-4o query. This is an estimate from assumed hardware, utilization and 500 output tokens, not a measurement. Epoch estimates about 2.5 Wh at 10,000 input tokens and about 40 Wh at 100,000.
- Sam Altman, June 2025 ("The Gentle Singularity"): "the average query uses about 0.34 watt-hours." No method, model or boundary given.
Side by side: our 2B classifier used 0.003 Wh, our 7B coding task used 0.20 to 0.38 Wh depending on retries, and our 27B used 0.62 Wh, GPU only. Google's active-machines-only figure (0.10 Wh) and full-facility figure (0.24 Wh) sit in between. In watt-hours per AI prompt, that's 0.003 to 0.62 Wh on our runs and 0.10 to 0.34 Wh in published figures. None of these share a workload, model or boundary, so read them as a range.
The rig runs in your VPC, or your engineers run it and we interpret the output.
Is a local GPU more efficient per request than a data center?
Not on our numbers so far: our one-request-at-a-time GPU used 0.20 Wh for a single 7B coding attempt, twice Google's active-machines-only median of 0.10 Wh for a production Gemini model, and our 27B at 0.62 Wh sits above Google's full-facility figure of 0.24 Wh.
The workloads aren't like-for-like: coding versus a median chat prompt, different models, different token counts. The likely reason for the gap is batching. A data center serves many requests per GPU at once; we measured batch size 1, one request at a time. Idle power on our card runs 12 to 17 W with the model loaded, and a busy data center spreads that same idle draw over far more requests than we do. We haven't measured batched serving on our own hardware yet, so treat batching as the likely explanation until we do.
What drives the cost of a request?
Five things moved the cost of a request in our data: output length, output format, active model size, retries and task difficulty.
- Output length. Reasoning and longer answers cost more because energy follows output tokens. Our hardest coding tasks pushed Qwen 3.8 27B to an average 1,849 output tokens and about 9,800 J per call, against 482 tokens and 2,221 J on ordinary tasks. See how retries change the token count.
- Output format. On Banking77 intent classification, the same 2B model used 24.3 J with a restricted label grammar versus 73.1 J with a JSON schema output, a 67% cut from the output format alone, with no model change.
- Active model size. Qwen3-Coder 30B is a mixture-of-experts model with a small fraction of parameters active per token; it used 448 J per coding task. Qwen 3.8 27B is dense and used 2,221 J on the same 164 tasks. On these two models, active parameters tracked the energy bill better than total size. See small vs large models on the same tasks.
- Retries. One attempt at a coding task cost 728 J. Adding a feedback retry ladder raised that to 1,367 J on average. Blind retries at the same token budget landed at 1,152 J.
- Task difficulty. The 34 hardest leftover coding tasks averaged about 9,800 J per Qwen 3.8 27B call, more than four times the 2,221 J on ordinary tasks, because harder problems make the model write more before it answers.
Full measurement disclosure for all of the above is on /results.
How to measure your own cost per request
To measure your own cost per request, log tokens per request for API calls and joules per request for self-hosted models, then divide by completed tasks:
- For an API: log input and output tokens per request from the provider's API response, then multiply each by that model's list price per token.
- For self-hosted models: sample GPU power during requests (amdgpu hwmon power1_average on AMD, nvidia-smi on NVIDIA) and subtract idle power measured with the model loaded but not serving.
- State your boundary. Say whether you're counting GPU only, the whole machine, or the whole facility, because that choice alone can more than double the number (Google's 0.10 vs 0.24 Wh).
- Divide by completed tasks, not calls. A request that fails and gets retried still cost energy and tokens; counting only successful calls understates the real cost per unit of work.
What hasn't been measured yet?
Batched serving is the biggest gap: everything above is batch size 1, and batching is the likely reason our per-request numbers run higher than Google's data-center figures. We also haven't measured full-machine or full-facility energy, only the GPU; hardware amortization per request, which is covered separately in the self-hosting cost breakdown; NVIDIA GPUs, since every local number here came from one AMD card; and long-context requests beyond the token counts in this post's table.
Frequently Asked Questions
How much does one LLM request cost?
Input tokens times the input price plus output tokens times the output price. 1,000 input and 500 output tokens cost $0.007 on Claude Sonnet 5 and $0.001 on Gemini 3.1 Flash-Lite. On our hard coding tasks, frontier calls ran $0.0003 to $0.012 each.
How much energy does an AI query use?
Google reports 0.24 Wh for a median Gemini text prompt, counting the whole data center. Epoch AI estimates about 0.3 Wh for a typical GPT-4o query. On our own GPU a 7B coding request used 0.20 Wh and a 27B used 0.62 Wh, GPU only.
How much electricity does a local LLM request cost?
Very little at the meter. A 7B coding task with retries used 0.38 Wh, about $0.00006 at $0.15/kWh. That excludes the hardware, which usually dominates; see the self-hosting cost breakdown.
Why do output tokens cost more than input tokens?
Output is generated one token at a time, so each output token takes a full pass through the model, while input tokens are processed together. Providers price that in, usually at about 5 times the input rate on the models above.
Do reasoning models use more energy per query?
Yes, because they write more tokens. On our hardest coding tasks a Qwen 3.8 27B call averaged 1,849 output tokens and about 9,800 J, against 482 tokens and 2,221 J on ordinary tasks.
Is self-hosting cheaper per request than an API?
Per request at the meter, usually yes. Whether it's cheaper overall depends on how busy you keep the GPU, since you pay for it whether it's serving or idle.
Methods
- Hardware. One AMD Radeon AI PRO R9700 (32GB) for every local request.
- Energy method. amdgpu hwmon power1_average sampled at about 5 Hz. Idle power measured over 60 s with the model loaded and subtracted from the per-request reading. GPU-only boundary, batch size 1.
- Classification boundary. ContractNLI and Banking77 rows come from a separate reference server, model-serving energy, batch size 1; see /results for the full disclosure fields.
- API costs. List price times tokens reported by each provider's API, run 2026-09-27.
- Electricity price. $0.15/kWh, a stated assumption, not a specific utility rate.
- Published-figure sources. Google Cloud blog and arXiv 2508.15734; epoch.ai/gradient-updates/how-much-energy-does-chatgpt-use; blog.samaltman.com/the-gentle-singularity.
- Data. Every measured row on this page: llm-cost-per-request-measured.csv.
- Dates. Local and classification runs published 2026-09-17 and 2026-09-27/28. API pricing and cost runs checked 2026-09-27 and 2026-09-28.