Qwen 3.8 27B and gpt-oss 20B tied for the best local LLM for coding in our test, each passing 84.8% of 507 hidden-test coding tasks, and gpt-oss did it with 45% less GPU energy and 13 GiB of memory.
All five models landed within 4 points of each other, a statistical tie.
Energy per task varied 18 times, from 137 J for a 7B to 2,524 J for GLM-4.7-Flash.
LeanLM (not affiliated with Google's LearnLM educational AI) measures what AI workloads cost to serve. Every number below comes from our own runs on a stated GPU.
Which local LLM is best for coding?
On pass rate alone, Qwen 3.8 27B and gpt-oss 20B tie at 84.8%, with gpt-oss faster and cheaper to run. We tested five models on 507 coding tasks (164 HumanEval+ and 343 MBPP+, scored on EvalPlus hidden tests (extra tests the model never sees)), each with up to 3 attempts per task, blind retry only when the code failed the task's visible tests (the examples in the task prompt), on one NVIDIA RTX PRO 6000 (96 GB) running Ollama:
| Model | Pass rate | HumanEval+ | MBPP+ | Time/task (median) | Energy/task (mean) | GPU memory (median) |
|---|---|---|---|---|---|---|
| Qwen 3.8 27B | 84.8% | 93.9% | 80.5% | 6.8 s | 1,543 J | 19.1 GiB |
| gpt-oss 20B | 84.8% | 94.5% | 80.2% | 4.3 s | 849 J | 13.0 GiB |
| GLM-4.7-Flash | 82.6% | 91.5% | 78.4% | 16.3 s | 2,524 J | 19.2 GiB |
| Qwen3-Coder 30B | 81.9% | 91.5% | 77.3% | 2.2 s | 260 J | 19.5 GiB |
| Qwen2.5-Coder 7B | 81.1% | 87.8% | 77.8% | 1.3 s | 137 J | 5.9 GiB |
Pass rate is on all 507 tasks combined; HumanEval+ and MBPP+ are the two benchmarks it's built from. Time is the median per task; energy is the mean per task (GPU power, idle-subtracted, summed over a task's attempts); GPU memory is the median nvidia-smi memory used during each model's run. gpt-oss vs Qwen 3.8 is derived: 45% less energy, about 38% less time, the same pass rate. One NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB), Ollama, default quantized tags (Q4 class (4-bit quantized)), num_ctx 16384, 2026-09-29. Full data: best-local-llm-for-coding-measured.csv and best-local-llm-for-coding-per-task.csv (every task).
Which model should you run?
Pick by what you're optimizing: pass rate and energy together point at gpt-oss, speed and cost point at Qwen3-Coder, and the smallest GPU points at the 7B. gpt-oss 20B gets the best pass rate at the lowest energy cost among the top two: 84.8%, tied with Qwen 3.8, at 849 J per task and 13.0 GiB. In Qwen3-Coder vs gpt-oss terms, Qwen3-Coder 30B is the fastest and cheapest model near the top: 81.9%, only 3 points off the leaders, at 2.2 s and 260 J per task, about a sixth of Qwen 3.8's energy (derived). Qwen2.5-Coder 7B is the pick for the smallest GPU: 81.1% at 5.9 GiB and 137 J per task, about 1/11 of Qwen 3.8's energy (derived), 3.7 points lower on pass rate.
Qwen 3.8 27B still makes sense when a problem is hard enough that the extra reasoning pays off: in an earlier test on a different GPU, it led on the hard leftovers other models couldn't solve, passing 63 of 102 attempts against Qwen3-Coder's 27. See that hard-leftover comparison for the setup. GLM-4.7-Flash reasons at length before answering, which is why it's both the slowest model here (16.3 s per task) and the most energy-hungry (2,524 J), without a pass-rate edge to show for it.
The rig runs in your VPC, or your engineers run it and we interpret the output.
Why is a tie the result?
None of the ten pairwise comparisons between these five models is a statistically significant difference in pass rate. We ran exact McNemar tests on paired per-task outcomes for all 10 pairs, Holm-adjusted for multiple comparisons: the smallest Holm-adjusted p-value is 0.34 (Qwen 3.8 vs the 7B, raw p = 0.034). Plainly: on these benchmarks the five models are statistically tied on pass rate. The differences that hold up are time, energy and memory.
Energy per passed task (total energy divided by passes, idle subtracted) tells the same story. The 7B is cheapest at 169 J per passed task, Qwen3-Coder next at 318 J, gpt-oss at 1,001 J, Qwen 3.8 at 1,819 J, and GLM highest at 3,054 J. A leaderboard that ranks by pass rate alone can't tell these five apart.
What does the retry loop add?
Most tasks were solved on the first attempt, and retries added a handful of passes: 29 for GLM-4.7-Flash, 21 for the 7B, 10 for gpt-oss, 9 for Qwen3-Coder and 3 for Qwen 3.8 out of 507. Qwen 3.8 and gpt-oss solved about 96 to 98% of tasks on attempt 1, Qwen3-Coder 92%, GLM 87% and the 7B 86%; mean attempts per task ranged 1.04 to 1.24 across the five models. Our full feedback retry ladder (a harness that feeds the failing test's error back to the model) exists for the 7B on a different GPU (an AMD Radeon AI PRO R9700): 144 of 164 HumanEval+ and 268 of 343 MBPP+. Blind retries here, which resend the original prompt with no error feedback, gave the 7B 144 and 267, and on an earlier run on that same AMD card, blind retries matched the ladder on HumanEval+ (143 vs 144) using 16% less energy. So blind retries are a fair stand-in for the ladder at 7B scale; we haven't measured whether feedback helps the larger models more. See our self-correction test and the fine-tune vs retry-ladder comparison for the harness mechanics.
Do the results repeat?
Mostly, within a few tasks. A first run the night before, on the same pod type and script, completed three of the five models before the pod failed; their pass counts on each benchmark differed from this run by 1 to 5 tasks: Qwen3-Coder went 260 then 265 on MBPP+, gpt-oss went 156 then 155 on HumanEval+, and the 7B went 142 then 144 on HumanEval+. Treat differences of a few tasks between runs as noise.
Best local LLM for coding on 16GB or 24GB VRAM
gpt-oss 20B is the best fit for a 16 GB VRAM card: it ran fully on the GPU at a 16k context using 13.0 GiB, which leaves about 3 GB free on a 16 GB card (derived from the measured figure; we didn't run it on a 16 GB card). Qwen2.5-Coder 7B fits comfortably too, at 5.9 GiB, small enough for an 8 GB card. For a local LLM for coding with 24GB VRAM, all five models fit at this context: Qwen 3.8 27B, GLM-4.7-Flash and Qwen3-Coder 30B ran at 19.1 to 19.5 GiB, which needs a 24 GB card at this context length (derived). Longer contexts add more memory on top of these figures; see how much VRAM a local LLM needs at different context lengths for that math.
How to run a local LLM for coding with Ollama
Pull the tag, set a context long enough for coding, give reasoning models room to answer, then check the output against tests. Here's the setup we used, for gpt-oss as an example:
ollama pull gpt-oss:20b
Set num_ctx to at least 16384 for coding tasks; a short context truncates longer prompts and files. Reasoning models like gpt-oss and Qwen 3.8 write their thinking before the answer, so give them room to finish: we ran with num_predict at 4096 in our tests. Then run the output against your tests, and retry on a failure; our retries were blind (the original prompt resent, no error passed back), and that added 3 to 29 passing tasks per model in our run.
What hasn't been tested yet?
Repository-level and multi-file tasks, the feedback retry ladder on the larger models, other quantizations, larger models, other runtimes and agent workflows are all untested here. This test covers single-function tasks in Ollama at Q4 quantization on one GPU: repository-level and multi-file work would have a model navigate and edit an existing codebase rather than write one function from a prompt; the feedback retry ladder is measured only on the 7B; other quantizations means Q8 and FP16; larger models means 70B and up, including Kimi and DeepSeek; other runtimes means vLLM and llama.cpp; agent workflows means calling tools and running commands mid-task.
Frequently Asked Questions
What is the best local LLM for coding?
In our test on 507 coding tasks, Qwen 3.8 27B and gpt-oss 20B tied at 84.8%. gpt-oss 20B used 45% less GPU energy per task and 13 GiB of memory, so it's the better pick unless you need Qwen 3.8's edge on harder problems.
What is the best local LLM for coding on 16GB VRAM?
gpt-oss 20B, which ran in 13.0 GiB at a 16k context and tied for the top pass rate. Qwen2.5-Coder 7B fits too, at 5.9 GiB, about 4 points lower.
What is the best local LLM for coding on 24GB VRAM?
Any of the five fits at a 16k context. gpt-oss 20B and Qwen 3.8 27B tied for the highest pass rate, and Qwen3-Coder 30B was the fastest of the three larger models at 2.2 s per task.
Is Qwen3-Coder better than gpt-oss for coding?
Not on pass rate: gpt-oss 20B passed 84.8% and Qwen3-Coder 30B 81.9%, a gap that isn't statistically significant. Qwen3-Coder was faster (2.2 s vs 4.3 s per task) and used less energy (260 J vs 849 J).
Which local coding model uses the least energy?
Qwen2.5-Coder 7B, at 137 J of GPU energy per task, then Qwen3-Coder 30B at 260 J. GLM-4.7-Flash used the most, 2,524 J, because it writes long reasoning before answering.
Do local LLMs need a retry loop for coding?
It helps smaller models most. Most tasks passed on the first attempt, and retrying on a failed visible test picked up the rest; for the 7B, simple blind retries matched our full feedback loop.
What's the best Ollama model for coding?
gpt-oss 20B, in our test: it tied Qwen 3.8 27B at 84.8% of 507 tasks with 45% less GPU energy and 13.0 GiB of memory. If speed matters most, Qwen3-Coder 30B was 3 points lower at 2.2 s per task.
Methods
- Hardware. RunPod, one NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB), 2026-09-29.
- Software. Ollama, default quantized tags (Q4 class): qwen3.8:latest, gpt-oss:20b, glm-4.7-flash:latest, qwen3-coder:30b, qwen2.5-coder:7b. num_ctx 16384, num_predict 4096, model default sampling.
- Tasks. 164 HumanEval+ and 343 MBPP+ (507 total), scored on EvalPlus hidden tests.
- Policy. Up to 3 attempts per task; retry only when the code failed the task's visible tests; retries resend the original prompt (blind, no error feedback).
- Energy method. GPU power streamed at 5 Hz; idle power (89 to 91 W with the model loaded) measured for 60 s and subtracted; energy per task is the sum over its attempts.
- Memory method. Median nvidia-smi memory used during each model's run (GiB).
- Statistics. Exact McNemar on paired per-task outcomes for all 10 pairs, Holm-adjusted for multiple comparisons.
- Known gap. GLM's power log has one 89-second gap from a network reconnect; energy across it was interpolated (about 1 to 2% of GLM's run).
- Repeatability run. A first run the night before (same pod type, same script) completed three models before the pod failed; used as a rough repeatability check.
- Data. Summary per model and benchmark: best-local-llm-for-coding-measured.csv. Every task: best-local-llm-for-coding-per-task.csv.
- Date. Measured 2026-09-29.