Skip to content

Measured energy per request, on your workload.

Find the waste in one workload

Loupe by LeanLM

Get more out of the GPUs you already have.

We measure what each AI request costs in tokens, seconds and joules on your workload, then show you where the waste is. On one task, changing only the output format cut energy per request by two thirds. Same model, same answers.

Your data doesn’t leave your environment. The method is published so you can check it.

Energy per request: JSON schema 73.1 joules, label grammar 24.3 joules, same model and task JSON schema 73.1 J Label grammar 24.3 J
Energy per request, same 2B model and task (3,072 cases). Only the output format changed.

⅓

The energy per request, same model, same answers. Only the output format changed.
intent task, 3,072 cases, one server

⅕

The energy of a dense 27B model, from a 30B mixture-of-experts model on the same 164 coding tasks
145 vs 152 tasks passed

51×

The energy reasoning cost for 2.6 points of accuracy
300-case banking test

You’re billed for tokens. You consume watts.

Those two numbers move on their own. A request that looks cheap on an invoice can be expensive in energy, and the gap is invisible from either side of a metered API. If you self-host, the watts are your bill. If you’re power-capped or queued for a GPU allocation, the watts are your ceiling. Either way, nobody is handing you the figure.

The reason it goes unmeasured is that it’s fiddly rather than hard. It needs a fixed server, a warmed idle baseline, a stated boundary, and the discipline to hold everything else constant. That’s the whole job.

Most of the energy wasn’t where we expected.

The obvious lever is model size, and it’s real: a 2B model used a twelfth of the energy of a 27B on the same intent task. The less obvious lever turned out to be larger in proportion. Holding the model fixed and changing only the output constraint moved per-request energy from 73.1 J to 24.3 J.

What changed Energy per request Change
27B model, label grammar 284.5 J baseline
2B model, JSON schema 73.1 J −74%
2B model, label grammar 24.3 J −91%

A JSON schema makes the model generate braces, quoted keys and whitespace before it reaches the answer. Every one of those is a decode step, and decode steps are where the joules go. Swapping the schema for a restricted label grammar removed two thirds of the remaining energy without touching the weights.

Measured at batch size 1 on one reference server. We haven’t yet measured whether the format gap survives production batching, and a restricted grammar only applies where the output space is bounded.

Since then we’ve run the same kind of measurement on coding. A 7B model with a retry loop passed 144 of 164 HumanEval+ tasks, matching the published single-attempt score of its 32B sibling. Retrying blind instead of feeding the error back came within one task of it (143 vs 144) with 16% less energy. Details: small vs large models and self-correction.

The six things a joule figure has to disclose.

An energy number without these is not checkable, and most published ones omit at least half. Ours are stated on every result, including the ones we can’t yet fill in.

Measurement tool and sampling rate

What read the power, and how often.

Boundary

GPU only, or the whole node. These differ by a lot and are frequently conflated.

Idle handling

Whether a warmed idle baseline was subtracted, and what it was.

Engine and concurrency

Serving engine, batch size, and the full server arguments.

Prefill and decode separately

Prompt tokens and generated tokens have very different energy profiles.

Attribution scheme

How a measured window was divided across requests.

Model-serving energy only. Retrieval, embedding and orchestration sit outside the boundary and are excluded rather than estimated. A model-serving figure is not a whole-system figure and we don’t present it as one.

What a measurement is for.

1

Finding the waste

The output-format result cost nothing to act on and removed two thirds of the energy on that task. On a saturated GPU that is serving capacity you get back without buying anything. Waste that size stays invisible until somebody puts a meter on it, and it tends not to be where people assume.

2

Picking the smallest model that clears your bar

Run your own tasks through small and large models and see pass rate and energy side by side. On our coding set, a 30B mixture-of-experts model passed 145 of 164 tasks at a fifth of the energy of a dense 27B that passed 152. Whether that trade is worth it depends on your bar, which is why we measure on your tasks.

3

Reporting it

Need the number in CO2e for a sustainability report? We convert measured energy with your region’s grid factor and state exactly what the boundary includes and leaves out.

What you get

Per-request table

Tokens, seconds and joules for your workload, with prefill and decode separated.

Six disclosures, filled in

So anyone on your team can check the number.

Changes ranked by savings

How much each change saves, and what it takes to make it.

The raw data

So your team can rerun and extend it.

One workload, fixed scope, fixed fee. Start with a short scoping call.

Frequently asked questions

Why does the output format change energy use?

A JSON schema makes the model generate its own syntax: braces, quoted keys, whitespace. Each of those is a decode step, and decode is where most inference energy goes. A restricted label grammar emits the answer and stops. On the task we measured, the same model produced the same answers for a third of the energy.

Does lower energy mean a lower bill?

If you self-host, yes, because you pay for the power and the GPU hours. On a metered API the relationship is indirect: you’re billed per token, and fewer output tokens is both fewer joules and a smaller invoice, but the provider’s margin sits between the two. We report the joules we measured and the token counts behind them, and leave the arithmetic on your contract to you.

How do you compare against published accuracy results?

Carefully, and often unfavourably. On Banking77 and ContractNLI there are task-specific encoder models under 350M parameters that score above the generative models we measured. That is itself an energy finding: running a general-purpose model where a specialised one would do has a measurable cost, and it is one of the things a measurement turns up.

Does our data leave our environment?

It doesn’t have to. The measurement rig runs inside your VPC, or your engineers run it and we interpret the output. That keeps the third-party risk assessment and the data-processing agreement off the critical path, which is usually where this kind of work actually stalls.

What hardware do you measure on?

A single reference server, stated on every result, with the GPU model and region named. We have replayed a run from AMD onto a rented NVIDIA GPU and matched 298 of 300 items; the two that didn’t match are reported rather than rounded away. Portability across hardware needs its own evidence and we don’t assert it.

What does a measurement cost?

Fixed scope and fixed fee for one workload, quoted after a short scoping call, with no ongoing dependency.

One workload. Fixed scope. Fixed fee.

Tell us where it runs and roughly how big. If measuring it won’t tell you anything useful, we’ll say so on the call.

Where does it run, and how big?

Thanks. We’ll reply within two business days.

Takes 20 seconds. No attachments. Security posture and GPU region come up on the call.

Get the next measurement by email.

One table at a time, no spam.

Done. You’ll get the next table when it lands.