Loupe by LeanLM
Get more out of the GPUs you already have.
We measure what each AI request costs in tokens, seconds and joules on your workload, then show you where the waste is. On one task, changing only the output format cut energy per request by two thirds. Same model, same answers.
Your data doesn’t leave your environment. The method is published so you can check it.
⅓
The energy per request, same model, same answers. Only the output format changed.
intent task, 3,072 cases, one server
⅕
The energy of a dense 27B model, from a 30B mixture-of-experts model on the same 164 coding tasks
145 vs 152 tasks passed
51×
The energy reasoning cost for 2.6 points of accuracy
300-case banking test
The problem
You’re billed for tokens. You consume watts.
Those two numbers move on their own. A request that looks cheap on an invoice can be expensive in energy, and the gap is invisible from either side of a metered API. If you self-host, the watts are your bill. If you’re power-capped or queued for a GPU allocation, the watts are your ceiling. Either way, nobody is handing you the figure.
The reason it goes unmeasured is that it’s fiddly rather than hard. It needs a fixed server, a warmed idle baseline, a stated boundary, and the discipline to hold everything else constant. That’s the whole job.
What we found
Most of the energy wasn’t where we expected.
The obvious lever is model size, and it’s real: a 2B model used a twelfth of the energy of a 27B on the same intent task. The less obvious lever turned out to be larger in proportion. Holding the model fixed and changing only the output constraint moved per-request energy from 73.1 J to 24.3 J.
| What changed | Energy per request | Change |
|---|---|---|
| 27B model, label grammar | 284.5 J | baseline |
| 2B model, JSON schema | 73.1 J | −74% |
| 2B model, label grammar | 24.3 J | −91% |
A JSON schema makes the model generate braces, quoted keys and whitespace before it reaches the answer. Every one of those is a decode step, and decode steps are where the joules go. Swapping the schema for a restricted label grammar removed two thirds of the remaining energy without touching the weights.
Measured at batch size 1 on one reference server. We haven’t yet measured whether the format gap survives production batching, and a restricted grammar only applies where the output space is bounded.
Since then we’ve run the same kind of measurement on coding. A 7B model with a retry loop passed 144 of 164 HumanEval+ tasks, matching the published single-attempt score of its 32B sibling. Retrying blind instead of feeding the error back came within one task of it (143 vs 144) with 16% less energy. Details: small vs large models and self-correction.
How we measure
The six things a joule figure has to disclose.
An energy number without these is not checkable, and most published ones omit at least half. Ours are stated on every result, including the ones we can’t yet fill in.
Measurement tool and sampling rate
What read the power, and how often.
Boundary
GPU only, or the whole node. These differ by a lot and are frequently conflated.
Idle handling
Whether a warmed idle baseline was subtracted, and what it was.
Engine and concurrency
Serving engine, batch size, and the full server arguments.
Prefill and decode separately
Prompt tokens and generated tokens have very different energy profiles.
Attribution scheme
How a measured window was divided across requests.
Model-serving energy only. Retrieval, embedding and orchestration sit outside the boundary and are excluded rather than estimated. A model-serving figure is not a whole-system figure and we don’t present it as one.
What it’s for
What a measurement is for.
Finding the waste
The output-format result cost nothing to act on and removed two thirds of the energy on that task. On a saturated GPU that is serving capacity you get back without buying anything. Waste that size stays invisible until somebody puts a meter on it, and it tends not to be where people assume.
Picking the smallest model that clears your bar
Run your own tasks through small and large models and see pass rate and energy side by side. On our coding set, a 30B mixture-of-experts model passed 145 of 164 tasks at a fifth of the energy of a dense 27B that passed 152. Whether that trade is worth it depends on your bar, which is why we measure on your tasks.
Reporting it
Need the number in CO2e for a sustainability report? We convert measured energy with your region’s grid factor and state exactly what the boundary includes and leaves out.
The deliverable
What you get
Per-request table
Tokens, seconds and joules for your workload, with prefill and decode separated.
Six disclosures, filled in
So anyone on your team can check the number.
Changes ranked by savings
How much each change saves, and what it takes to make it.
The raw data
So your team can rerun and extend it.
One workload, fixed scope, fixed fee. Start with a short scoping call.
Questions
Frequently asked questions
Why does the output format change energy use?
A JSON schema makes the model generate its own syntax: braces, quoted keys, whitespace. Each of those is a decode step, and decode is where most inference energy goes. A restricted label grammar emits the answer and stops. On the task we measured, the same model produced the same answers for a third of the energy.
Does lower energy mean a lower bill?
If you self-host, yes, because you pay for the power and the GPU hours. On a metered API the relationship is indirect: you’re billed per token, and fewer output tokens is both fewer joules and a smaller invoice, but the provider’s margin sits between the two. We report the joules we measured and the token counts behind them, and leave the arithmetic on your contract to you.
How do you compare against published accuracy results?
Carefully, and often unfavourably. On Banking77 and ContractNLI there are task-specific encoder models under 350M parameters that score above the generative models we measured. That is itself an energy finding: running a general-purpose model where a specialised one would do has a measurable cost, and it is one of the things a measurement turns up.
Does our data leave our environment?
It doesn’t have to. The measurement rig runs inside your VPC, or your engineers run it and we interpret the output. That keeps the third-party risk assessment and the data-processing agreement off the critical path, which is usually where this kind of work actually stalls.
What hardware do you measure on?
A single reference server, stated on every result, with the GPU model and region named. We have replayed a run from AMD onto a rented NVIDIA GPU and matched 298 of 300 items; the two that didn’t match are reported rather than rounded away. Portability across hardware needs its own evidence and we don’t assert it.
What does a measurement cost?
Fixed scope and fixed fee for one workload, quoted after a short scoping call, with no ongoing dependency.
Scope a measurement
One workload. Fixed scope. Fixed fee.
Tell us where it runs and roughly how big. If measuring it won’t tell you anything useful, we’ll say so on the call.
Takes 20 seconds. No attachments. Security posture and GPU region come up on the call.
Not ready for a call?
Get the next measurement by email.
One table at a time, no spam.