Loupe by LeanLM
Get more out of the GPUs you already have.
We measure what each AI request costs in tokens, seconds, joules and dollars on your workload, then show you where the waste is. On one task, changing only the output format cut energy per request by two thirds. Same model, same answers.
Your data doesn’t leave your environment. The method is published so you can check it.
⅓
The energy per request, same model, same answers. Only the output format changed.
intent task, 3,072 cases, one server
⅕
The energy of Qwen 3.8 27B (dense), from Qwen3-Coder 30B (mixture-of-experts) on the same 164 coding tasks
145 vs 152 tasks passed
51×
The energy reasoning cost for 2.6 points of accuracy
300-case banking test
The problem
You’re billed for tokens. You consume watts.
Those two numbers move on their own. A request that looks cheap on an invoice can be expensive in energy, and the gap is invisible from either side of a metered API. If you self-host, the watts are your bill. If you’re power-capped or queued for a GPU allocation, the watts are your ceiling. Either way, nobody is handing you the figure.
It goes unmeasured because it’s fiddly. You fix the server, subtract a warmed idle baseline, state the boundary, and change one thing at a time.
On a metered API
No GPUs? The same waste shows up in your token bill.
On a hosted API you can’t meter the watts, so we measure what you’re billed for: input and output tokens, latency, and dollars per completed task, per feature. The format change below cut output from about 37 tokens to about 7 per request, and time per request from 1.15 s to 0.25 s. We measured that on an open model; on yours we measure it on your traffic. Changes come back ranked by dollars saved: output format, prompt caching, batch pricing, model choice.
What we found
Most of the energy wasn’t where we expected.
Model size matters. On the same intent task, a 2B model with a label grammar used a twelfth of the energy of a 27B running with no output constraint. The output format mattered more in proportion: holding the 2B fixed and changing only the output constraint moved per-request energy from 73.1 J to 24.3 J.
| What changed | Energy per request | Change |
|---|---|---|
| 27B model, no output constraint | 284.5 J | baseline |
| 2B model, JSON schema | 73.1 J | −74% |
| 2B model, label grammar | 24.3 J | −91% |
A JSON schema makes the model generate braces, quoted keys and whitespace before it reaches the answer. Every one of those is a decode step, and decode steps are where the joules go. Swapping the schema for a restricted label grammar removed two thirds of the remaining energy without touching the weights.
The table above is batch size 1 on one reference server. The format gap holds under batching, narrower: on vLLM with Qwen2.5-1.5B, constrained choice used 48% less energy per request than a JSON schema one request at a time, and 36% less at 32 concurrent. Batching itself cut energy per request 11 to 14 times (the batched run). A restricted grammar only applies where the output space is bounded.
Same 2B model, 0.25 s per request instead of 1.15 s. At $2 per GPU-hour and batch size 1, that’s about $140 per million requests instead of $640.
Since then we’ve run the same kind of measurement on coding. A 7B model with a retry loop passed 144 of 164 HumanEval+ tasks, matching the published single-attempt score of its 32B sibling, Qwen2.5-Coder-32B. Retrying blind, without feeding the error back, passed 143 and used 16% less energy. Both retry styles cost more than one attempt: 1,152 J per task blind and 1,367 J with feedback, against 728 J for a single try. Details: small vs large models and self-correction.
How we measure
The six things a joule figure has to disclose.
An energy number without these is not checkable, and most published ones omit at least half. Ours are stated on every result, including the ones we can’t yet fill in.
Measurement tool and sampling rate
What read the power, and how often.
Boundary
GPU only, or the whole node. These differ a lot and often get mixed up.
Idle handling
Whether a warmed idle baseline was subtracted, and what it was.
Engine and concurrency
Serving engine, batch size, and the full server arguments.
Prefill and decode separately
Prompt tokens and generated tokens have very different energy profiles.
Attribution scheme
How a measured window was divided across requests.
Model-serving energy only. Retrieval, embedding and orchestration sit outside the boundary and aren’t in the number, not even as an estimate.
What it’s for
What a measurement is for.
Finding the waste
The output-format change cost nothing to make and removed two thirds of the energy on that task. On a saturated GPU, that’s serving capacity back without buying more hardware.
Picking the smallest model that clears your bar
Run your own tasks through small and large models and see pass rate and energy side by side. On our coding set, Qwen3-Coder 30B (mixture-of-experts) passed 145 of 164 tasks at a fifth of the energy of Qwen 3.8 27B (dense), which passed 152. Whether 7 fewer passes is worth that depends on your bar, so we run it on your tasks.
Reporting it
Need the number in CO2e for a sustainability report? We convert measured energy with your region’s grid factor and state exactly what the boundary includes and leaves out.
The deliverable
What you get
Per-request table
Tokens, seconds, joules and dollars for your workload, with prefill and decode separated. Dollars come from your GPU-hour rate or your API price list.
Six disclosures, filled in
So anyone on your team can check the number.
Changes ranked by savings
How much each change saves, and what it takes to make it.
The raw data
So your team can rerun and extend it.
One workload, fixed scope, fixed fee. Start with a short scoping call.
Getting started
What we need from you
A workload
A sample of real requests or a test set, plus the quality bar it has to clear.
Access
A GPU node or VPC slot for the rig, or an engineer who runs it with us. For an API workload, a key with a spending cap.
Time
One engineer who knows the workload, for the kickoff and the readout.
Security and paperwork
We sign your MSA and DPA. We don’t have SOC 2 yet, so we answer your security questionnaire instead, and we finish vendor onboarding before we touch anything.
Questions
Frequently asked questions
Why does the output format change energy use?
A JSON schema makes the model generate its own syntax: braces, quoted keys, whitespace. Each of those is a decode step, and decode is where most inference energy goes. A restricted label grammar emits the answer and stops. On the task we measured, the same model produced the same answers for a third of the energy.
Does lower energy mean a lower bill?
If you self-host, yes, because you pay for the power and the GPU hours. On a metered API the relationship is indirect: you’re billed per token, and fewer output tokens is both fewer joules and a smaller invoice, but the provider’s margin sits between the two. We report the joules and the token counts behind them, and price the tokens at your contract rate.
How do you compare against published accuracy results?
Carefully, and often unfavourably. On Banking77 and ContractNLI there are task-specific encoder models under 350M parameters that score above the generative models we measured. That’s an energy finding too: running a general-purpose model where a specialised one would do has a measurable cost, and a measurement shows it.
Does our data leave our environment?
It doesn’t have to. The measurement rig runs inside your VPC, or your engineers run it and we interpret the output. That keeps the third-party risk assessment and the data-processing agreement off the critical path, which is usually where this kind of work actually stalls.
What hardware do you measure on?
A single reference server, stated on every result. We replayed a run from AMD onto a rented NVIDIA GPU and 298 of 300 items matched. The two that didn’t are reported with the run. Results hold across the hardware we’ve tested.
What does a measurement cost?
$5,000 to $15,000 for one workload, depending on how many models and configurations we compare. About two weeks from access to readout. Fixed fee, agreed before we start. Our first engagements run below that range, at a design-partner rate.
You’re pre-launch. Why go first?
The method is published and every figure we cite traces to a run record. Our first engagements run at a design-partner rate, in exchange for a reference call or a short case study you approve.
Scope a measurement
One workload. Fixed scope. Fixed fee.
$5,000 to $15,000 for one workload, about two weeks. Our first engagements run below that, at a design-partner rate. Tell us where it runs and roughly how big. If measuring it won’t tell you anything useful, we’ll say so on the call.
Takes 20 seconds. No attachments.
Not ready for a call?
Get the next measurement by email.
One table at a time, no spam.