Qwen 3.8 27B fits on one 32 GB AMD R9700 at Q4_K_M and peaks near 24 GB of VRAM.
In one attempt it passed 92.7% of HumanEval+ (152 of 164) at a median 6.7 seconds and 79 tokens per second.
On harder tasks it beat Qwen3-Coder 30B clearly (63 vs 27 of 102); on HumanEval as a whole the two are close (152 vs 145, p = 0.17).
Loupe by LeanLM (not affiliated with Google's LearnLM educational AI) measures what AI workloads cost to serve. Tokens, seconds and joules per task are measured, not estimated.
How much VRAM does Qwen 3.8 27B need?
About 17 GB for the weights at Q4_K_M (a 4-bit quantization), and up to 24 GB once you allocate room for a long context. On our AMD Radeon AI PRO R9700 (32 GB), Ollama's Vulkan backend loaded qwen3.8:latest, the 27.3B-parameter dense build of Qwen 3.8, and the model used 22.7 GB with a 128k-token context allocated during the hard-task run and peaked at 23.8 GB during the full HumanEval+ run. Either way, it leaves headroom on a 32 GB card; A 16 GB card can't hold the weights at this quant; we haven't tested smaller cards with a shorter context.
Cold load from disk took about 5 minutes on this machine, which has slow storage; a faster disk will load faster. Idle power with the model loaded was 12 W. Qwen3-Coder 30B, the 30.5B mixture-of-experts build we ran alongside it, needs more room for its larger weight file and peaked higher, at 25.8 GB, even though it's a MoE model with fewer active parameters per token.
How fast is Qwen 3.8 27B on an AMD R9700?
79 tokens per second decode, measured from Ollama's own eval_count and eval_duration on the full HumanEval+ run (72 to 85 tok/s from the 10th to 90th percentile). End to end, including prompt processing, that's 71 tokens per second, and the median task took 6.7 seconds (mean 8.2 s, max 32 s, 0 timeouts) at a median 482 output tokens.
We ran with speculative multi-token-prediction (MTP) decoding on, draft_num_predict set to 4: the model drafts several tokens ahead and verifies them in one pass. That's the likely reason it decodes faster than a plain one-token-at-a-time estimate for this card; we didn't measure it with MTP off. Published single-R9700 numbers for this model that we've seen range around 25 to 32 tokens per second, with one report showing 24.9 tok/s and 31.9 tok/s with MTP on. Setups differ: backend, quant and speculative-decoding settings all move the number, so treat any single figure, including ours, as specific to its configuration. Ours is Ollama on Vulkan, Q4_K_M, with MTP on.
Qwen3-Coder 30B, which doesn't think before answering, decoded faster: 139 tokens per second median, 129 tok/s end to end, at a median 1.8 seconds per task.
How good is Qwen 3.8 27B at coding?
92.7% of HumanEval+ in one attempt, no retries, run on 2026-09-28. Full results, both models, 164 tasks each:
Download the raw per-task results: Qwen 3.8 27B vs Qwen3-Coder 30B, HumanEval+ (CSV) and, for the hard-task comparison below, the fallback-model bake-off (CSV).
| Model | HumanEval+ | Base | Median s/task | Decode tok/s | J/task | Peak VRAM |
|---|---|---|---|---|---|---|
| Qwen 3.8 27B | 152/164 (92.7%) | 161/164 (98.2%) | 6.7 s | 79 | 2,221 (2,396/pass) | 23.8 GB |
| Qwen3-Coder 30B | 145/164 (88.4%) | 151/164 (92.1%) | 1.8 s | 139 | 448 (507/pass) | 25.8 GB |
One attempt per task, 164 tasks, hidden EvalPlus tests, Ollama defaults, one AMD Radeon AI PRO R9700, run 2026-09-28.
Paired, Qwen 3.8 solved 13 tasks Qwen3-Coder missed, and Qwen3-Coder solved 6 that Qwen 3.8 missed. Exact McNemar p = 0.17 (a paired test on the tasks where only one model passed): no detectable difference on the full HumanEval+ set. The gap shows up on harder problems. We took the 34 coding tasks a Qwen2.5-Coder-7B retry loop couldn't solve in an earlier test (10 HumanEval, 24 MBPP+), ran each model 3 attempts per task against hidden tests (102 attempts total), and Qwen 3.8 27B passed 63 of 102, solving at least one attempt on 24 of the 34 tasks, at a median 16 seconds per call (mean 38 s; 3 of 102 hit the 300-second timeout) and about 9,800 J per call. Qwen3-Coder 30B passed 27 of 102, solving 10 of the 34, at a median 3.5 seconds and about 980 J per call. Holm-adjusted p = 0.0024, a significant gap. For reference on the same subset, gpt-6-astra passed 84 of 102 and gpt-6-luna 70.
Qwen 3 Coder vs Qwen 3.8: which should you run?
Qwen 3.8 for accuracy on hard tasks: it reasons before it answers, and that's where the 63 vs 27 gap comes from. Qwen3-Coder for speed and energy on routine functions where the two tie: it's about 4x faster and uses about 5x less energy per task, and on the full HumanEval+ set the two models aren't statistically different (152 vs 145, p = 0.17). If your workload is mostly straightforward functions with the occasional hard one, route by difficulty rather than picking one model for everything. We haven't measured routing itself here, just the two models on the same tasks, so treat that as guidance from the numbers above, not a tested claim about other workloads.
The rig runs in your VPC, or your engineers run it and we interpret the output.
Does it matter where you run the same weights?
Yes, a lot. The same Qwen 3.8 27B weights on Cloudflare Workers AI passed 45 of 102 on the hard-task subset, against 63 of 102 running locally, and 50 of the 102 calls on Workers AI errored or timed out. Same model, different host, different result. If you're calling a hosted reasoning model, set max_tokens explicitly and use a long timeout: a request that doesn't set max_tokens can cut a reasoning model off mid-thought before it ever writes an answer, which is the same max_tokens pitfall we measured on Workers AI reasoning models in an earlier test.
What harness should you run Qwen 3.8 in?
The only harness we've measured it in: a test-driven retry harness with Qwen 3.8 27B as the local fallback model behind a fine-tuned 7B. In one diagnostic run, that combination took the system to 90.2% HumanEval+ and 98.8% Base. That number is a diagnostic, not a clean benchmark: the 7B's training data specifically targeted HumanEval failures, so it isn't a clean held-out number. The retry-loop mechanics are in our 7B retry-ladder measurement, and the case for retrying blind instead of feeding the model its error is in our self-correction test. We haven't tested third-party agent harnesses, Claude Code, OpenCode and similar, with Qwen 3.8 yet.
What hasn't been tested yet?
Other quants: we ran Q4_K_M only, not Q8, FP8 or BF16. Repository-level tasks: everything above is single-function HumanEval and MBPP+ style problems, not a multi-file codebase with its own build. Long context: we allocated 128k tokens of context during the hard-task run, but the prompts themselves were short, so this isn't a test of long-context quality or speed. Other GPUs and other backends: we ran one AMD R9700 on Ollama's Vulkan backend, not vLLM or ROCm, and not an NVIDIA card.
Frequently Asked Questions
How much VRAM does Qwen 3.8 27B need?
At Q4_K_M the weights are about 17 GB; with a 128k context allocated it used 22.7 to 23.8 GB on a 32 GB AMD R9700.
Does Qwen 3.8 27B run on an AMD Radeon AI PRO R9700?
Yes, fully on the GPU with Ollama's Vulkan backend: 79 tokens per second decode and a median 6.7 seconds per HumanEval task.
Which Qwen 3.8 quant should I use?
We measured Q4_K_M, which passed 92.7% of HumanEval+ in one attempt and fits a 32 GB card with room for a 128k context. We haven't measured other quants yet.
Is Qwen 3 Coder or Qwen 3.8 better for coding?
On HumanEval+ they're close (92.7% vs 88.4%, not a significant gap). On the hardest tasks Qwen 3.8 passed 63 of 102 attempts vs 27. Qwen3-Coder is about 4x faster and uses about 5x less energy per task.
What is the best coding harness for Qwen 3.8?
The one we've measured is a test-driven retry harness with Qwen 3.8 as the local fallback model. We haven't tested third-party agent harnesses with it yet.
Methods
- Hardware. One AMD Radeon AI PRO R9700 (32GB).
- Backend. Ollama, Vulkan.
- Models. qwen3.8:latest (Qwen 3.8 27B, 27.3B dense, Q4_K_M, 262,144-token native context, thinking on, temperature 1, top_k 20, top_p 0.95, speculative multi-token-prediction decoding on with draft_num_predict 4) and qwen3-coder:30b (Qwen3-Coder 30B, 30.5B mixture of experts, Q4_K_M, no thinking), both on Ollama defaults otherwise.
- Task sets. Full HumanEval+ (164 tasks, one attempt each, no retries). A hard-task subset: the 34 coding tasks (10 HumanEval, 24 MBPP+) a Qwen2.5-Coder-7B retry loop couldn't solve in an earlier test, 3 attempts per task (102 attempts per model).
- Scoring. EvalPlus hidden tests.
- Energy. amdgpu hwmon power1_average sampled at 5 Hz, idle power subtracted, GPU-only boundary.
- Stats. Exact McNemar test on the paired full HumanEval+ outcomes; Holm correction for the hard-subset comparison.
- Data. Full per-task results: qwen-3-8-27b-humaneval-plus-r9700.csv, llm-self-correction-fallback-models.csv.