One coding request on a 7B model running on our own GPU took 238 tokens, 5.65 seconds and 0.38 Wh with retries, about $0.00006 of electricity.

A hard coding request to a frontier API cost $0.0003 to $0.012 at list price, because the models spent 141 to 3,080 output tokens on it.

Energy per request ranged from 0.003 Wh for a 2B classifier to 0.62 Wh for a 27B coding answer; Google's published median Gemini prompt is 0.24 Wh.

Loupe by LeanLM (not affiliated with Google's LearnLM educational AI) measures what AI workloads cost to serve. Every number below comes from our own runs or a named published source.

How do you calculate LLM cost per request?

LLM cost per request is input tokens times the input price, plus output tokens times the output price. Both prices are quoted per million tokens, so divide by 1,000,000 before multiplying.

Worked example: 1,000 input tokens and 500 output tokens on Claude Sonnet 5 ($2/$10 per million). Input costs 1,000 × $2 / 1,000,000 = $0.002. Output costs 500 × $10 / 1,000,000 = $0.005. Total: $0.007. The same request on Gemini 3.1 Flash-Lite ($0.25/$1.50 per million) costs 1,000 × $0.25 / 1,000,000 + 500 × $1.50 / 1,000,000 = $0.00025 + $0.00075 = $0.001, about a seventh of the Sonnet 5 price. That's your cost per API call.

Output tokens usually dominate the bill. They're priced at 5 to 6 times the input rate on the models above, and a model that thinks before answering pays for that thinking in output tokens.

Tokens per request aren't fixed by the task either. On the same hard coding tasks, list-price API calls ran from $0.0003 to $0.01172 because the models spent very different amounts of output on the same problem: gpt-6-astra wrote 141 tokens per call, gpt-6-luna wrote 557, and gemini-3.8-flash wrote 3,080. Price per token alone doesn't tell you price per request; you need the token count too.

What one request costs, measured

Measured, one request cost between $0.0000004 of electricity for a 2B classifier and $0.01172 at list price for the most verbose frontier answer:

Request Model Output tokens Seconds Joules Wh Cost
Coding task, one attemptQwen2.5-Coder-7B7280.20$0.00003 electricity
Coding task, feedback retriesQwen2.5-Coder-7B238 (mean)5.65 (mean)1,367 (mean)0.38 (mean)$0.00006 electricity
Coding task, blind retriesQwen2.5-Coder-7B182 (mean)1,1520.32$0.00005 electricity
Coding task, one attemptQwen3-Coder 30B (MoE)221 (median)1.8 (median)4480.12$0.00002 electricity
Coding task, one attemptQwen 3.8 27B482 (median)6.7 (median)2,2210.62$0.00009 electricity
ContractNLI classification2B model9.00.0025$0.0000004 electricity
Banking77 classification2B model24.30.0067$0.0000010 electricity
Hard coding taskgpt-6-luna (API)557 (mean)5.9 (mean)$0.0003 list price
Hard coding taskgpt-6-astra (API)141 (mean)4.1 (mean)$0.00918 list price
Hard coding taskgemini-3.8-flash (API)3,080 (mean)10.6 (mean)$0.01172 list price

Local rows: GPU-only energy, idle power subtracted, one AMD Radeon AI PRO R9700, one request at a time (batch size 1). Classification rows: model-serving energy, batch size 1, one reference server, published 2026-09-17 on /results. API rows: list price × tokens reported by the provider API, run 2026-09-27; provider-side energy isn't visible to us. Electricity priced at $0.15/kWh; excludes hardware purchase, cooling and the rest of the machine. See the hardware side of the cost. Joules are the mean per task or call; 1 Wh = 3,600 J. Full data: llm-cost-per-request-measured.csv.

How much energy does an AI query use?

A typical published figure for one text query is 0.24 to 0.34 Wh; our own requests ranged from 0.003 Wh to 0.62 Wh depending on model size and task. Energy per ChatGPT query, per Epoch AI's estimate, is about 0.3 Wh.

Bar chart of energy per request in watt-hours: 2B classifier 0.003, Qwen3-Coder 30B coding task 0.12, 7B coding task 0.20, Google median Gemini prompt 0.10 active machines only and 0.24 full data center, 7B with retries 0.38, Qwen 3.8 27B 0.62
Energy per request in watt-hours. Ours: GPU only, one AMD R9700, batch size 1. Google: published August 2025.

Side by side: our 2B classifier used 0.003 Wh, our 7B coding task used 0.20 to 0.38 Wh depending on retries, and our 27B used 0.62 Wh, GPU only. Google's active-machines-only figure (0.10 Wh) and full-facility figure (0.24 Wh) sit in between. In watt-hours per AI prompt, that's 0.003 to 0.62 Wh on our runs and 0.10 to 0.34 Wh in published figures. None of these share a workload, model or boundary, so read them as a range.

The rig runs in your VPC, or your engineers run it and we interpret the output.

Is a local GPU more efficient per request than a data center?

Not on our numbers so far: our one-request-at-a-time GPU used 0.20 Wh for a single 7B coding attempt, twice Google's active-machines-only median of 0.10 Wh for a production Gemini model, and our 27B at 0.62 Wh sits above Google's full-facility figure of 0.24 Wh.

The workloads aren't like-for-like: coding versus a median chat prompt, different models, different token counts. The likely reason for the gap is batching. A data center serves many requests per GPU at once; we measured batch size 1, one request at a time. Idle power on our card runs 12 to 17 W with the model loaded, and a busy data center spreads that same idle draw over far more requests than we do. We haven't measured batched serving on our own hardware yet, so treat batching as the likely explanation until we do.

What drives the cost of a request?

Five things moved the cost of a request in our data: output length, output format, active model size, retries and task difficulty.

Full measurement disclosure for all of the above is on /results.

How to measure your own cost per request

To measure your own cost per request, log tokens per request for API calls and joules per request for self-hosted models, then divide by completed tasks:

What hasn't been measured yet?

Batched serving is the biggest gap: everything above is batch size 1, and batching is the likely reason our per-request numbers run higher than Google's data-center figures. We also haven't measured full-machine or full-facility energy, only the GPU; hardware amortization per request, which is covered separately in the self-hosting cost breakdown; NVIDIA GPUs, since every local number here came from one AMD card; and long-context requests beyond the token counts in this post's table.

Frequently Asked Questions

How much does one LLM request cost?

Input tokens times the input price plus output tokens times the output price. 1,000 input and 500 output tokens cost $0.007 on Claude Sonnet 5 and $0.001 on Gemini 3.1 Flash-Lite. On our hard coding tasks, frontier calls ran $0.0003 to $0.012 each.

How much energy does an AI query use?

Google reports 0.24 Wh for a median Gemini text prompt, counting the whole data center. Epoch AI estimates about 0.3 Wh for a typical GPT-4o query. On our own GPU a 7B coding request used 0.20 Wh and a 27B used 0.62 Wh, GPU only.

How much electricity does a local LLM request cost?

Very little at the meter. A 7B coding task with retries used 0.38 Wh, about $0.00006 at $0.15/kWh. That excludes the hardware, which usually dominates; see the self-hosting cost breakdown.

Why do output tokens cost more than input tokens?

Output is generated one token at a time, so each output token takes a full pass through the model, while input tokens are processed together. Providers price that in, usually at about 5 times the input rate on the models above.

Do reasoning models use more energy per query?

Yes, because they write more tokens. On our hardest coding tasks a Qwen 3.8 27B call averaged 1,849 output tokens and about 9,800 J, against 482 tokens and 2,221 J on ordinary tasks.

Is self-hosting cheaper per request than an API?

Per request at the meter, usually yes. Whether it's cheaper overall depends on how busy you keep the GPU, since you pay for it whether it's serving or idle.

Methods