A small language model (SLM) is usually one under about 10B parameters, small enough to run on one GPU.
On the same 164 coding tasks, a 7B with a retry loop passed 144, a 27B passed 152 in one attempt, and a 30B mixture-of-experts model that uses about 3B parameters per token passed 145 at a fifth of the 27B's energy.
Small models clear the bar when the output can be checked and a larger model catches what they miss.
Loupe by LeanLM (not affiliated with Google's LearnLM educational AI) measures what AI workloads cost to serve. Every token, second and joule on this page comes from our own runs.
What is a small language model?
A small language model is one under about 10B parameters, or one that runs on a single GPU.
There's no agreed cutoff, and sources draw the line in different places:
- IBM describes SLMs as ranging from "a few million to a few billion" parameters.
- Hugging Face's SLM overview uses roughly 1 million to 10 billion.
- The survey "Small Language Models: Survey, Measurements, and Insights" (arXiv 2409.15790) covers models from 100M to 5B.
- NVIDIA's "Small Language Models are the Future of Agentic AI" (Belcak et al., arXiv 2506.02153, June 2025) skips a parameter count and defines an SLM by what it does: it fits on a common consumer device and answers one user's requests with practical latency. NVIDIA says most models below 10B qualified as of 2025.
Our working line
That's the line we use on this page. Mixture-of-experts models muddy this: Qwen3-Coder 30B has 30.5B total parameters but only about 3.3B active per token, so it downloads like a large model and runs like a small one.
SLM vs LLM on the same coding tasks
Same 164 HumanEval+ (164 Python function tasks scored on extra hidden tests from EvalPlus) tasks, one AMD Radeon AI PRO R9700 (32 GB), GPU-only energy with idle power subtracted:
| Model | Setup | Passed | Pass rate | J/task | Time/task | Tokens/task |
|---|---|---|---|---|---|---|
| Qwen 3.8 27B (dense) | One attempt | 152/164 | 92.7% | 2,221 (median) | 6.7 s (median) | 482 (median) |
| Qwen3-Coder 30B (MoE, ~3.3B active) | One attempt | 145/164 | 88.4% | 448 (median) | 1.8 s (median) | 221 (median) |
| Qwen2.5-Coder-7B + retry loop | Retry loop | 144/164 | 87.8% | 1,367 (mean) | 5.65 s (mean) | 238 (mean) |
Runs 2026-09-27 and 2026-09-28. All numbers on this page: slm-vs-llm-small-vs-large-measured.csv. The same 7B with blind resampling instead of feedback scored 143/164 at 1,152 J and 182 tokens per task, close to the feedback version at less energy. The 27B vs 30B MoE gap on the full set isn't significant (exact McNemar p = 0.17, a paired test; the difference is within what chance would produce). Per-task results for both models: qwen-3-8-27b-humaneval-plus-r9700.csv. Qwen 3.8 and Qwen3-Coder times and tokens are medians; the 7B row is a mean.
For a published reference point outside our own runs, Qwen2.5-Coder-32B-Instruct scores 87.2% on HumanEval+ in a single attempt on the EvalPlus leaderboard. Our 7B with retries matches that number. It isn't a clean comparison: theirs is one attempt, and ours retries against the visible tests.
The plain reading: on this workload the energy bill tracked active parameters and output length, not the size printed on the model card. The 30B mixture-of-experts model was the cheapest of the three to run, despite being the largest download.
Does a retry loop help small models more?
On our small set, the loop helped the 4B most. On 35 held-out MBPP+ (basic Python programming tasks, same hidden-test scoring) tasks, one attempt vs. with a retry loop (run 2026-09-24):
| Model | One attempt | With retry loop |
|---|---|---|
| Qwen2.5-0.5B-Instruct | 37.1% (13/35) | 48.6% (17/35) |
| Qwen3.5-4B | 42.9% (15/35) | 62.9% (22/35) |
| Qwen2.5-Coder-7B-Instruct | 71.4% (25/35) | 77.1% (27/35) |
35 tasks is a small set; one task is worth about 3 points. The same 35 tasks were used while choosing the loop's rules, so treat the loop numbers as a best case. Energy wasn't measured in this run.
The loop added the most for the 4B (+7 tasks, 15 to 22). The 0.5B still fails most tasks even with the loop.
The rig runs in your VPC, or your engineers run it and we interpret the output.
Where do small models fall short?
On the hardest coding tasks, the local 27B trailed the frontier APIs by a wide margin: 63 vs 84 of 102 attempts.
We took the 34 coding tasks (10 HumanEval, 24 MBPP+) that the 7B retry loop never solved, ran three attempts each against hidden tests (102 attempts per model), and checked how large and frontier models did on the same leftovers (runs 2026-09-27):
| Model | Passed | Notes |
|---|---|---|
| gpt-6-astra (OpenAI API) | 84/102 | $0.00918 per call at list price |
| gemini-3.8-flash (Google API) | 82/102 | $0.01172 per call at list price |
| gpt-6-luna (OpenAI API) | 70/102 | $0.0003 per call at list price |
| Qwen 3.8 27B (local) | 63/102 | ~9,800 J, ~38 s per call avg |
| Qwen3-Coder 30B (local) | 27/102 | ~980 J per call |
The 7B retry loop scores 0 on this set by construction: it's the set of tasks the 7B never solved. See the full 36-model table for the complete fallback bake-off.
The local 30B MoE trailed further still, at 27 of 102.
Small models for classification
Coding isn't the only workload. From our energy-per-request measurements (published 2026-09-17, a 2B and a 27B model from one family, one reference server, batch size 1, model-serving energy):
- ContractNLI (1,763 cases): the 2B scored 84.7% at 9.0 J per request; the 27B scored 78.5% at 649.1 J.
- Banking77 (3,072 cases): the 2B scored 92.0% at 24.3 J per request with a restricted label grammar; the 27B scored 93.9% at 284.5 J.
ContractNLI's origin paper reports a 110M-parameter BERT-base model at 83.8% and a 335M variant at 87.5%, both task-fine-tuned in a different configuration, so it's not a like-for-like number against our runs. The point still holds: for narrow classification, the right small model may not be generative at all. See the full results page for the disclosure fields on every run, or cost per request for the dollar and joule math behind these numbers.
When should you use a small language model?
When to use a small language model: when the output can be checked automatically (by tests, a schema or a fixed label set) and the task is narrow. Put a verifier and a larger fallback model behind it for the leftovers (see how we route to a fallback model and how routing works more generally).
Use a large model first when you can't check the output automatically, when tasks are open-ended, or when the leftover rate is high enough that escalation dominates the bill. In our coding runs the small tier handled most of the work and a fallback caught the rest: the 7B loop missed 20 of 164 HumanEval+ tasks, and the best frontier fallback passed 84 of 102 attempts on the hard leftovers.
Are small language models the future of agentic AI?
NVIDIA's paper (arXiv 2506.02153) argues SLMs are "sufficiently powerful, inherently more suitable, and necessarily more economical" for many agentic calls. Our data backs the economics on checkable, narrow tasks and also shows the limit: on the hard leftovers, the local 27B trailed the frontier APIs, 63 vs. 84 of 102. We haven't measured agent workloads yet. On what we have measured: small first where the output is checkable, with a frontier fallback for the tasks that need one.
What hasn't been tested yet?
Repository-level and multi-file tasks: everything above is single-function problems, not a real codebase with its own build. Agent workloads: we measured single-turn generation, not multi-step tool use. Other model families: everything here is Qwen; we haven't run the same comparison on Llama, Mistral or a different lineage. Batched serving: all our energy numbers are batch size 1, and batching changes the joules-per-request math. Other GPUs: everything ran on one AMD Radeon AI PRO R9700.
Frequently Asked Questions
What is the difference between an SLM and an LLM?
Size and where it runs. SLMs are usually under about 10B parameters and fit on one GPU or device; LLMs are larger and usually served from a data center. There's no agreed cutoff: IBM, Hugging Face and NVIDIA draw the line differently.
Are small language models as accurate as LLMs?
On narrow, checkable tasks they can be close. On 164 HumanEval+ coding tasks a 7B with a retry loop passed 144 and a 27B passed 152. On the 34 hardest leftovers, the best frontier API passed 84 of 102 attempts and a local 27B passed 63.
How much cheaper is it to run an SLM than an LLM?
How much cheaper an SLM is depends on active parameters and output length more than the label. On one GPU, a 30B mixture-of-experts model used 448 J per coding task, a 7B with retries 1,367 J and a dense 27B 2,221 J. On contract classification a 2B used 9.0 J per request against 649.1 J for a 27B.
Can a fine-tuned small model beat a large one?
On narrow tasks, yes. On ContractNLI our 2B scored 84.7% vs a 27B's 78.5%, and the task's origin paper reports a 110M fine-tuned encoder at 83.8%.
Can small language models run on a laptop?
Models under a few billion parameters usually can. We measured everything here on one 32 GB AMD Radeon AI PRO R9700 desktop GPU, not a laptop.
When should I use a small language model instead of an LLM?
When the output can be checked automatically and the task is narrow. Put a verifier and a larger fallback model behind it so the misses get caught.
Methods
- Hardware. One AMD Radeon AI PRO R9700 (32GB) for every local run.
- Backends. Ollama, Vulkan, for Qwen 3.8 27B and Qwen3-Coder 30B; the Qwen2.5-Coder-7B, Qwen2.5-0.5B-Instruct and Qwen3.5-4B retry-loop runs used the same R9700.
- Scoring. EvalPlus hidden tests for every HumanEval+ and MBPP+ number on this page.
- Energy. amdgpu hwmon power1_average sampled at about 5 Hz, idle power subtracted, GPU-only boundary. Not measured for the MBPP+ size-ladder run.
- Stats. Exact McNemar test on the paired full HumanEval+ outcomes.
- Classification boundary. ContractNLI and Banking77 numbers are from a separate reference server, model-serving energy, batch size 1; see /results for the full disclosure fields.
- Data. Every number on this page: slm-vs-llm-small-vs-large-measured.csv.
- Dates. Coding runs 2026-09-24, 2026-09-27 and 2026-09-28. Classification numbers published 2026-09-17.