Self-correction without an external check doesn't help. With tests to check against, code repair does.
At 7B, the retry does the work and the error message adds nothing measurable: blind retries passed 143 of 164 HumanEval+ tasks vs 144 for test-feedback repair (p = 1.0), using 24% fewer tokens and 16% less GPU energy.
When retries run out, send the task once to a reasoning model: gpt-6-astra solved 84 of 102 attempts on the hardest leftovers; a local 27B model solved 63.
Loupe by LeanLM (not affiliated with Google's LearnLM educational AI) measures what AI workloads cost to serve. Tokens, seconds and joules per task are measured, not estimated.
Does LLM self-correction work?
It depends on whether the model has something outside itself to check against. Without external feedback, self-correction doesn't reliably improve reasoning, and it can make reasoning worse: Huang et al. (arXiv 2310.01798) found LLMs struggle to fix their own reasoning mistakes when they're the only judge. Olausson et al. (arXiv 2306.09896) tested the code case directly: self-repair helps only when the feedback is good, and the gains shrink once you count the extra samples it costs to generate that feedback. Kamoi et al. (arXiv 2406.01297), surveying the field, conclude that reliable self-correction needs an external verifier: tests, a compiler, something the model isn't grading itself. Arimbur (arXiv 2604.10508) shows what that verifier buys: with tests in place, iterative repair adds 4.9 to 17.1 points on HumanEval across seven models, and most of that lands in the first two rounds.
Tests are what make code the case where self-correction pays.
Should a coding agent feed the error back or just retry?
Retry. At 7B, showing the model its error didn't buy any accuracy, and a blind retry costs less. Loupe ran a matched-budget control on Qwen2.5-Coder-7B-Instruct, one AMD Radeon AI PRO R9700: same number of calls in both arms, but a blind retry gets the original prompt only, no failure text. Blind retries passed 143 of 164 HumanEval+ tasks; feedback repair passed 144 (exact McNemar p = 1.0, a paired test where 1.0 means no detectable difference). Blind retries used 182 tokens per task against 238, and 1,152 J against 1,367 (GPU-only, idle subtracted): 24% fewer tokens and 16% less energy for the same result. Rerun at three fixed seeds, the feedback ladder scored 143, 144 and 142, so its one-task edge over blind retries is inside run-to-run noise.
Verma (arXiv 2607.26117) found the same pattern on MBPP+ from 1.5B to 7B, and has an explanation for why: shown its own failed attempt, a small model writes a program nearly identical to the one that just failed in 33 to 68% of retries. A blind retry, starting fresh, repeats itself only 2 to 14% of the time. At this size the failure text added nothing measurable. What helps is a second try that isn't anchored to the first. The full ladder, training runs and tables are in our 7B retry-ladder measurement.
Retries hold up off HumanEval too. On 343 MBPP+ tasks never used to design the ladder, the ladder passed 268 against 248 for a single attempt: 20 tasks gained, 0 lost, p = 1.9e-06. Of the tasks it recovered on the visible tests, 74% also passed the hidden tests. That's the ladder compared with a single attempt, not with blind retries; MBPP+ didn't run a blind arm.
How does the retry harness know a task failed?
The code runs against the task's visible tests in a sandboxed subprocess with a timeout. A failed assertion, an exception or a timeout triggers a retry. The hidden HumanEval+ tests only run after the loop finishes, for scoring; the harness never sees them. A blind retry uses the exact same check to decide whether to retry again, its prompt just leaves the error out. The step-by-step version, with the pre-check and tier budgets, is in the retry-ladder post.
How many retries before escalating?
Cap them early. The ladder's retries added 6 to 8 tasks over a single attempt, at 1,367 J per task against 728 J: about 15 kJ for each extra task solved. Arimbur (above) finds most of the gain in the first two rounds. Practical rule: cap cheap retries, then escalate once.
Which fallback model should a coding agent escalate to?
Send the leftovers to a model that reasons before it answers. Setup: the 34 tasks the 7B's retry loop couldn't solve (10 HumanEval, 24 MBPP+), sent to 36 models, 3 attempts each, scored on hidden tests.
Download the raw results: fallback-model bake-off (CSV) and blind vs. feedback retry data (CSV).
| Model | Where it runs | Hidden-test passes (of 102) | Time per attempt | Cost per attempt |
|---|---|---|---|---|
| gpt-6-astra | OpenAI | 84 | 4.1 s | $0.0092 |
| gemini-3.8-flash (thinking high) | 82 | 10.6 s | $0.0117 | |
| gpt-6-sol | OpenAI | 76 | 4.1 s | $0.0030 |
| GLM-5.2 | Cloudflare Workers AI | 73 | 39.8 s | $0.0124 |
| Kimi K2.6 | Cloudflare Workers AI | 72 | 78.6 s | $0.0170 |
| GLM-5.3 Flash | Cloudflare Workers AI | 71 | 78.5 s | $0.0010 |
| DeepSeek V4 Flash | Cloudflare Workers AI | 71 | 65.0 s | $0.0048 |
| gpt-6-luna | OpenAI | 70 | 5.9 s | $0.0003 |
| gpt-oss-120b | Cloudflare Workers AI | 69 | 28.2 s | $0.0015 |
| Kimi K2.7 Code | Cloudflare Workers AI | 67 | 45.9 s | $0.0085 |
| Gemma 4 26B | Cloudflare Workers AI | 67 | 92.9 s | $0.0011 |
| GLM-5.3 | Cloudflare Workers AI | 66 | 22.9 s | $0.0054 |
| DeepSeek V4 Pro | Cloudflare Workers AI | 65 | 96.6 s | $0.0082 |
| Nemotron 3 120B | Cloudflare Workers AI | 65 | 31.1 s | $0.0038 |
| Qwen 3.8 27B | Local, one AMD R9700 | 63 | 38.1 s | ~9,800 J (electricity ≈ $0.0004) |
| Qwen3 30B-A3B | Cloudflare Workers AI | 56 | 43.1 s | $0.0022 |
| gpt-oss-20b | Cloudflare Workers AI | 53 | 25.0 s | $0.0011 |
| Qwen 3.8 27B | Cloudflare Workers AI | 45 | 83.9 s | $0.0031 |
| GLM-4.7 Flash | Cloudflare Workers AI | 43 | 104.9 s | $0.0017 |
| Qwen2.5-Coder-32B | Cloudflare Workers AI | 31 | 6.7 s | $0.0003 |
| Qwen3-Coder 30B | Local, one AMD R9700 | 27 | 5.4 s | ~980 J |
34 tasks × 3 attempts = 102 per model, hidden EvalPlus tests. Timeouts (300 s) and API errors count as failures. List prices checked September 28, 2026 from OpenAI, Google and Cloudflare pricing pages. Local rows are electricity only at $0.15/kWh; hardware is excluded. 14 more models, all at 27 of 102 or below, are in the CSV. A 36th, Llama 3.2 11B Vision, couldn't be run: it requires a license agreement we hadn't signed.
Every model that scored 65 or above reasons before it answers. The two Qwen coder models, which don't, finished at 31 and 27. No model is statistically ahead of gpt-6-astra at 34 tasks: every Holm-adjusted comparison came back p ≥ 0.5.
Speed separates the top group more than accuracy does. OpenAI's models answer in 4 to 6 seconds; the Workers AI reasoning models take 23 to 97 seconds, and some time out. Same weights, different host, different result: Qwen 3.8 27B solved 63 of 102 running locally and 45 of 102 on Workers AI, where 50 of 102 calls errored or timed out.
Our picks: default to gpt-6-astra (about $0.56 per 1,000 coding tasks, if 6% of them escalate as on HumanEval; MBPP escalated 7%); gpt-6-luna if you're budget-constrained; a fully local Qwen 3.8 27B if you want zero external calls; on Cloudflare, gpt-oss-120b or GLM-5.2.
Set max_tokens on Workers AI
When a request doesn't set max_tokens, many Workers AI models stop at 256 output tokens. A reasoning model spends that whole budget thinking and never writes any code: gpt-oss-120b scored 0 of 102 that way, and 69 of 102 once max_tokens was set to 16,384.
The rig runs in your VPC, or your engineers run it and we interpret the output.
What hasn't been tested yet?
Single-function tasks only, HumanEval and MBPP. Repository-level work, with its own build and fixtures, is next. One main model at 7B; Verma sees the anchoring penalty shrink as models get stronger, so a larger main model may get more out of feedback than this one did. And the escalation set is small, 34 tasks, so differences under about 6 tasks between fallback models aren't detectable.
Frequently Asked Questions
Does self-repair help small code models?
With tests, yes, but the retry does the work. At 7B, blind retries matched test-feedback repair (143 vs 144 of 164 on HumanEval+) with 24% fewer tokens.
Is LLM self-correction worth the extra tokens?
For code with tests, a few retries are: they added 6 to 8 HumanEval+ tasks for a 7B model. Showing the model its error didn't add accuracy, so blind retries are the cheaper way to get it.
Which model should a coding agent escalate to when it's stuck?
A model that reasons before answering. On the 34 tasks a 7B couldn't solve, gpt-6-astra passed 84 of 102 attempts, gpt-6-luna 70, and a local Qwen 3.8 27B 63.
Does self-correction use more energy?
Yes. The 7B retry ladder used 1,367 J per task vs 728 J for one attempt. Blind retries cut that to 1,152 J at the same accuracy.
How many retries before escalating to a bigger model?
Cap cheap retries at a few rounds; most gains come in the first two. Then send the task once to a fallback model.
Methods
- Model. Qwen2.5-Coder-7B-Instruct, one AMD Radeon AI PRO R9700 (32GB).
- Benchmarks. HumanEval+ (164 tasks) and MBPP+ (343 tasks, held out from ladder design), scored on EvalPlus hidden tests.
- Blind control. Same call budget as the feedback ladder; retries get the original prompt and no failure text.
- Energy. amdgpu hwmon power1_average sampled at 5 Hz, idle power (16 to 17 W) subtracted, GPU-only boundary.
- Fallback bake-off. The 34 tasks the stock ladder exhausted (10 HumanEval, 24 MBPP+), routed by visible-test failure only, 3 independent attempts per model, each a single call with no feedback.
- Stats. Exact McNemar test on paired outcomes, Holm correction for multiple comparisons.
- Data. Full per-model and per-seed results: fallback-models.csv, blind-vs-feedback.csv.