Model Comparison
Every fully tested model on this hardware, sorted however you like. Click any column header to sort. Performance numbers come from llama-bench with pp512, tg128, full GPU offload, batch 2048, and flash attention. The standard thread count is 32. Puzzle is the measured exception at 10 threads. Quality scores come from automated test suites and human evaluation. Partial results from earlier evaluations are listed separately below.
Performance
| Model | Quant | Size | RADV pp | RADV tg | AMDVLK pp | AMDVLK tg |
|---|---|---|---|---|---|---|
| Gemma-4-E2B | UD-Q4_K_XL | 2.9 GiB | 3382 | 109 | 954 | 99 |
| Granite-4.1-3B | UD-Q4_K_XL | 2.0 GiB | 2278 | 88.7 | 652 | 88.1 |
| Gemma-4-E4B | UD-Q4_K_XL | 4.7 GiB | 1828 | 59 | 491 | 57 |
| GPT-OSS-20B-Derestricted | MXFP4 | 11 GiB | 1405 | 77 | 1380 | 77 |
| Qwen3.5-4B | Q8_0 | 4 GiB | 1375 | 37.9 | 510 | 40.0 |
| Gemma-4-26B-A4B | Q8_0 | 25 GiB | 1303 | 47.7 | 790 | 47.1 |
| Qwen3-30B-Instruct-2507 | UD-Q4_K_XL | 16.5 GiB | 1143 | 92 | 936 | 93 |
| Nemotron-3-Nano-30B-A3B | UD-Q4_K_XL | 21.3 GiB | 1106 | 68 | 776 | 65 |
| Nemotron-3-Nano-Omni-30B-A3B-Reasoning | UD-Q4_K_XL | 22.8 GiB | 1097 | 61.1 | 777 | 58.7 |
| Nex-N2-mini | s-batman Q8_0 | 34.4 GiB | 1083 | 54.5 | - | - |
| Qwen3.6-35B-A3B | Unsloth UD-Q4_K_XL | 20.8 GiB | 1029 | 59.6 | 615 | 57.6 |
| Ornith-1.5-35B-A3B New | Official Q8_0 | 35.20 GiB | 804.19 | 49.93 | n/a | n/a |
| Ornith-1.0-35B | Q8_0 | 34.36 GiB | 1027 | 53.2 | 725 | 52.5 |
| Agents-A1 New | Official Q8_0 | 34.37 GiB | 1010.49 | 53.79 | 723.54 | 53.48 |
| Qwen3.5-35B-A3B | Unsloth UD-Q4_K_XL | 21 GiB | 1017 | 59.8 | 686 | 60.4 |
| GLM-4.7-Flash | UD-Q4_K_XL | 16.3 GiB | 990 | 73 | 529 | 70 |
| Qwen3.5-9B | UD-Q4_K_XL | 5.6 GiB | 972 | 35.7 | 289 | 35.2 |
| Ornith-1.5-9B New | Official Q8_0 | 8.87 GiB | 830.81 | 23.75 | n/a | n/a |
| Nemotron-Cascade-2-30B-A3B | Q8_0 | 31 GiB | 968 | 54 | n/a | n/a |
| Granite-4.1-8B | UD-Q4_K_XL | 5.1 GiB | 936 | 38.6 | 265 | 38.7 |
| Kimi-Linear-48B-A3B | Q8_0 | 48.6 GiB | 746 | 52 | 570 | 53 |
| Gemma-4-12B | Unsloth UD-Q8_K_XL | 13.6 GiB | 716 | 14.0 | 199 | 14.0 |
| Ministral-3-14B | UD-Q4_K_XL | 7.8 GiB | 696 | 25 | 173 | 25 |
| GPT-OSS-120B | MXFP4 | 59 GiB | 596 | 56 | 661 | 53 |
| Qwen3-Coder-Next-80B | MXFP4 | 41 GiB | 586 | 40 | 462 | 43 |
| Magistral-Small-2509 | UD-Q4_K_XL | 13.5 GiB | 389 | 15 | 94 | 15 |
| Devstral-Small-2-24B | UD-Q4_K_XL | 13.5 GiB | 382 | 15 | 94 | 15 |
| Mistral-Small-4-119B | Unsloth UD-Q4_K_XL | 69 GiB | 363 | 40.3 | 313 | 39.2 |
| Qwen3.6-27B | Unsloth UD-Q4_K_XL | 16.4 GiB | 322 | 12.0 | 86 | 12.0 |
| Qwen3.5-27B | UD-Q4_K_XL | 16 GiB | 310 | 12.1 | 86 | 11.9 |
| Qwen3.5-122B-A10B | Unsloth UD-Q4_K_XL | 72 GiB | 287 | 22.4 | 197 | 21.9 |
| Granite-4.1-30B | UD-Q4_K_XL | 16.5 GiB | 275 | 11.8 | 71.2 | 11.9 |
| Puzzle 75B-A9B New | YanissAmz Q4_K_M | 48.05 GiB | 231.87 | 2.69 | 36.03 | 2.68 |
| Laguna S 2.1 | Poolside Q4_K_M | 70.01 GiB | 263 | 21.63 | 198 | 20.52 |
| Gemma-4-31B | Unsloth UD-Q4_K_XL (Apr 11) | 17.5 GiB | 261 | 11.1 | 70.8 | 11.1 |
| MiniMax-M2.7 | Unsloth UD-IQ4_XS | 101 GiB | 181 | 27.0 | 179 | 24.8 |
| MiniMax-M2.5 | Unsloth UD-Q3_K_XL | 94 GiB | 179 | 22 | 164 | 32 |
| Qwen3.8-27B New | Official Q8_0 | 26.63 GiB | 236.7 | 7.5 | - | - |
| Muse-Glimmer-30B | Unsloth UD-Q8_K_XL | 30.07 GiB | 234.1 | 7.3 | - | - |
| DeepSeek-V4-Flash-0731 | Unsloth UD-IQ3_XXS | 97.05 GiB | 114.7 | 12.4 | - | - |
| Mistral-Medium-3.5-128B | Unsloth UD-Q4_K_XL | 70.5 GiB | 62 | 2.9 | 17 | 2.9 |
Quality
The Tooling column is new and tracks automation reliability out of 65. It is scored separately and does not count toward the Combined total. Models marked with a hyphen have not run it yet. Scores here are for each model's reference quant. The 31B row keeps its Unsloth quant numbers as the stable reference. Google's QAT Q4_0 matches it within noise on quality and pulls ahead on tooling, artifact size, and speed. Its full-suite breakdown and the complete comparison are on the quantization page.
Qwen3.8-27B also keeps its official Q8_0 row as the reference. The Unsloth UD-Q4_K_XL follow-up scored 221/255 on the six medium-reasoning technical benchmarks against 219/255 for Q8, then completed all eight comparable technical suites in 2 hours 32 minutes instead of 12 hours 58 minutes. I have not assigned Q4 a Combined score because its nonthinking writing profile has not run. The benchmark table and the controlled speed comparison are on the quantization page.
| Model | Writing /30 | LRU /10 | FastAPI /8 | LeetCode /59 | Polyglot /65 | Postgres /57 | Cassandra /56 | Tooling /65 | Combined /285 |
|---|---|---|---|---|---|---|---|---|---|
| Qwen3.8-27B New | 23 | 10 | 8 | 59 | 65 | 42 | 35 | 62.5 | 242 |
| Muse-Glimmer-30B | 26 | 10 | 8 | 59 | 64 | 36 | 37 | 64.5 | 240 |
| DeepSeek-V4-Flash-0731 | 23 | 10 | 8 | 59 | 63 | 40 | 43 | 53.5 | 246 |
| Gemma-4-31B | 27 | 10 | 8 | 59 | 44 | 51 | 39 | 54.5 | 238 |
| Ornith-1.0-35B New | 28 | 10 | 8 | 59 | 57 | 46 | 28 | 61 | 236 |
| Ornith-1.5-35B-A3B New | 22 | 10 | 8 | 59 | 64 | 32 | 31 | 62.5 | 226 |
| Agents-A1 New | 22 | 10 | 8 | 59 | 54 | 40 | 31 | 59.5 | 224 |
| Gemma-4-12B | 28 | 10 | 8 | 59 | 42 | 47 | 38 | 54 | 232 |
| Gemma-4-26B-A4B | 28 | 10 | 8 | 59 | 22 | 49 | 39 | 59 | 215 |
| Qwen3.6-27B | 30 | 10 | 8 | 59 | 30 | 44 | 32 | 60.5 | 213 |
| Mistral-Medium-3.5-128B | 30 | 7 | 8 | 59 | 30 | 39 | 34 | - | 207 |
| Qwen3.6-35B-A3B | 29 | 10 | 8 | 59 | 35 | 31 | 33 | 53.5 | 205 |
| Puzzle 75B-A9B New | 27 | 10 | 8 | 59 | 31 | 35 | 34 | 47.5 | 204 |
| Laguna S 2.1 | 21 | 10 | 8 | 59 | 43 | 29 | 36 | 57.5 | 206 |
| Qwen3.5-122B-A10B | 29 | 10 | 8 | 59 | 13 | 36 | 37 | - | 192 |
| Qwen3.5-35B-A3B | 28 | 10 | 8 | 59 | 17 | 32 | 38 | - | 192 |
| Kimi-Linear-48B-A3B | 30 | 10 | 8 | 59 | 22 | 30 | 31 | - | 190 |
| Qwen3.5-27B | 25 | 10 | 8 | 59 | 21 | 34 | 29 | - | 186 |
| MiniMax-M2.5 | 26 | 10 | 7 | 59 | 13 | 40 | 30 | - | 185 |
| Nex-N2-mini | 28 | 10 | 8 | 53 | 17 | 38 | 29 | 57 | 183 |
| GPT-OSS-120B | 20 | 10 | 8 | 59 | 22 | 40 | 31 | - | 190 |
| MiniMax-M2.7 | 28 | 10 | 8 | 59 | 8 | 38 | 28 | - | 179 |
| Ornith-1.5-9B New | 15 | 10 | 7 | 57 | 24 | 35 | 27 | 59 | 175 |
| Gemma-4-E4B | 23 | 7 | 8 | 59 | 13 | 35 | 34 | - | 179 |
| Mistral-Small-4-119B | 21 | 10 | 7 | 59 | 34 | 27 | 26 | - | 184 |
| Qwen3-30B-Instruct-2507 | 30 | 10 | 2 | 59 | 13 | 27 | 31 | - | 172 |
| Granite-4.1-30B | 28 | 10 | 2 | 59 | 10 | 32 | 31 | - | 172 |
| Qwen3-Coder-Next-80B | 26 | 10 | 2 | 59 | 9 | 33 | 32 | - | 171 |
| Devstral-Small-2-24B | 27 | 10 | 2 | 59 | 11 | 29 | 31 | - | 169 |
| GPT-OSS-20B-Derestricted | 13 | 10 | 8 | 59 | 14 | 37 | 23 | - | 164 |
| Gemma-4-E2B | 18 | 7 | 8 | 59 | 10 | 27 | 27 | - | 156 |
| Nemotron-3-Nano-Omni-30B-A3B-Reasoning | 24 | 8 | 2 | 59 | 6 | 31 | 18 | - | 148 |
| Ministral-3-14B | 26 | 2 | 2 | 59 | 20 | 23 | 18 | - | 150 |
| Qwen3.5-9B | 20 | 10 | 0 | 51 | 14 | 28 | 23 | - | 146 |
| Granite-4.1-8B | 18 | 4 | 4 | 59 | 10 | 26 | 24 | - | 145 |
| Nemotron-Cascade-2-30B-A3B | 18 | 10 | 8 | 59 | 1 | 22 | 21 | - | 139 |
| Qwen3.5-4B | 16 | 9 | 8 | 54 | 3 | 17 | 16 | - | 123 |
| Nemotron-3-Nano-30B-A3B | 20 | 4 | 0 | 46 | 7 | 16 | 16 | - | 109 |
| Magistral-Small-2509 | 20 | 0 | 8 | 30 | 9 | 12 | 35 | - | 114 |
| Granite-4.1-3B | 17 | 0 | 0 | 53 | 4 | 11 | 13 | - | 98 |
| GLM-4.7-Flash | 14 | 0 | 0 | 16 | 0 | 23 | 27 | - | 80 |
Key Findings
- I prefer Ornith-1.5-35B-A3B among the two new models. I tested the verified official 35.20 GiB Q8_0 with thinking enabled and the published sampler. It reaches 226/285 Combined with 10/10 LRU, 8/8 FastAPI, 59/59 LeetCode, and 64/65 Polyglot. Tooling adds 62.5/65 with every call machine usable, and calibration is 22/22. Database review and diagnosis are strong while procedural authoring remains weak. PostgreSQL reaches 32/57 and Cassandra reaches 31/56. RADV gives 804.19 prompt tokens per second and 49.93 generation tokens per second. The full suite takes 129.7 minutes. Several correct coding answers still run for four to twelve minutes, which makes the model more verbose than I want from a routine local assistant. Creative writing reaches 22/30. Its separate prose sample ranks second in a blind comparison against the 9B and a human reference, with better control than the smaller model but less character depth than the reference.
- Ornith-1.5-9B reaches 175/285 Combined from the verified official 8.87 GiB Q8_0. I used its precise coding sampler for technical work and its separate general sampler for writing. Bounded tasks remain competitive at 10/10 LRU, 57/59 LeetCode, 59/65 Tooling, and 22/22 calibration. Longer tasks reveal poor stopping behavior. Polyglot falls to 24/65 and takes 170.9 minutes. Several Polyglot and database prompts fill most of the 64K context without producing a usable final answer. The technical suite takes 558 minutes even though generation runs at 23.75 tokens per second. Creative writing reaches 15/30, with broad continuity and pacing problems. The prose sample ranks third in the same blind comparison.
- I started Qwen3.8-27B on release day with the official 26.63 GiB Q8_0 target and 2.95 GiB Q8_0 MTP drafter. It finishes at 242/285 Combined, second behind DeepSeek-V4-Flash at 246 and one point ahead of Gemma-4-31B QAT at 241. LRU, FastAPI, LeetCode, and Polyglot are all perfect, with the first 65/65 Polyglot score on the board. PostgreSQL reaches 42/57 and Cassandra reaches 35/56. Both database runs sweep review, optimization, and diagnosis, then lose most of their points on procedural authoring. Tooling reaches 62.5/65 at 98.5 percent usable reliability and calibration reaches 22/22. Medium reasoning gives the technical work plenty of room, though the full suite takes 14.83 hours and the creative writing request reaches the 60 minute cap without a final answer. I retained that failed mode qualification and scored the official nonthinking run instead. It finishes 2083 words in 269 seconds for 23/30. Dialogue and sensory control are strong, while over-explanation, overly symmetrical character work, and a continuity miss cost points. The separate prose run finishes 2443 words in 312 seconds with all seven beats and exact mechanics. It ranks fourth in the five-sample comparison, held back by safe phrasing, familiar imagery, and a setting-consistency error. The fixed server control gives MTP a clean comparison. Generation rises from 5.10 to 11.87 tokens per second for a 2.33x gain at 60.84 percent acceptance. Two retained rescoring artifacts document the only score corrections. The Polyglot Apache fixture had the wrong expected status count, and the calibration matcher missed a typographic apostrophe. The generated answers were unchanged.
- Muse Glimmer 30B is the first Meta release I have benched since the Llama line went quiet, and it lands at 240/285 Combined the day after release. That puts it third, behind DeepSeek-V4-Flash at 246 and Gemma-4-31B QAT at 241, a strong showing for a dense 30B against a 284B mixture of experts. It takes two board bests on the way. Polyglot reaches 64/65, one ahead of DeepSeek, and Tooling reaches 64.5/65 with every call machine usable and the prompt injection resisted. The base weights are bfloat16 at 59.6 GB with no quantization config. Meta's own GGUF repo ships 19.65 GB and 16.76 GB variants, while the tested Unsloth Q8_K_XL is 30.07 GiB and preserves finer precision. Across both database benchmarks it scores a perfect 10/10 on review and optimization and a perfect 6/6 on diagnosis, then falls to 14/31 and 13/30 on procedural authoring. It understands databases better than it writes them. Creative writing reaches 26/30 with board-leading sensory work and dialogue, then loses points for labeling emotional beats instead of dramatizing them. The prose test ranks ahead of DeepSeek-V4-Flash and Agents-A1 and leads on voice risk, but it finishes last on genre fulfillment and stops at 1512 words against a 2400 to 3400 target. Plain generation runs 7.25 tokens per second at 234 prompt processing on RADV. Meta's DFlash drafter raises generation to 13.53 tokens per second for a 1.89x gain at 59 percent draft acceptance. The uncapped Tooling run remains comparable because the old ceiling never bound on an earlier row.
- The Ruby polyglot challenge was broken for its entire history and I have now fixed it, which moved six models. Across 104 recorded runs not one model had ever scored a single point on the Sinatra webhook task, which I had been reading as a hard challenge. It was two harness bugs stacked. Sinatra 4.1 added host authorization, which answers 403 and the text "Host not permitted" to the test client's default host header before any route code runs, so every request failed no matter what the model wrote. Underneath that, the result parser only recognised minitest 4 wording, "N tests", while minitest 5 emits "N runs", so Ruby results parsed as zero even when the suite passed. I confirmed both by writing a correct reference app to the published spec and watching it score zero. The fix lives in the verifier rather than the prompt, because the task is about webhook logic and not about Sinatra deployment hardening. I then re-ran every stored Ruby submission through the corrected harness without re-querying any model, and 29 runs gained points. This was not a uniform lift. DeepSeek-V4-Flash, Gemma-4-12B and Mistral-Small-4 had been writing perfect Ruby all along, Ornith was at 9 of 10, and Agents-A1 genuinely fails on its own merits. A second challenge was wrong in the same spirit. The FastAPI rate limiter resets module state between tests, but its skip list ignored any variable whose name began with an underscore, which is the ordinary way to mark module-private state in Python. A model that called its store _store had rate limit counts leak between tests and lost four points for the naming choice alone. The identical implementation scores 5 with an underscore and 9 without. Nineteen more runs gained points once that was fixed. Between the two bugs, polyglot has been scored out of roughly 55 while presented as out of 65. Original result files are untouched and both corrections are recorded alongside them.
- DeepSeek-V4-Flash-0731 is a 284B mixture of experts with roughly 13B active, and it takes the top combined score at 246/285. I ran it on Unsloth's published card, temp 1.0, top-p 1.0, and min-p 0.0, with top-k and both penalties pinned to no-ops because llama.cpp otherwise applies its own top-k of 40 and quietly contradicts a top-p of 1.0. That is the same practice I follow elsewhere, running each model on its documented settings where the vendor publishes them. Cassandra 43/56 is a clear board best, four ahead of the previous 39, with Tier 2 perfect at 10/10 and Tier 3 perfect at 6/6. Polyglot 63/65 is also a board best by a wide margin, and chasing it is what uncovered two broken challenges described further down. LRU, FastAPI, LeetCode, and the hallucination calibration set are all perfect. Creative writing lands at 23/30. Dialogue is exceptionally strong, while the sample loses points for weak genre fulfillment and over-explained imagery. Speed is the practical problem. I measured 114.7 prompt processing and 12.4 generation on RADV, which is poor for a 13B-active model, and the cause is missing kernels rather than memory pressure. The server disables the Lightning Indexer and three fused hash-clustering ops at load because Vulkan has no implementation for them, and cutting context from 65536 to 40960 cleared most of the swap while moving prompt processing only from 111 to 114. ROCm is not available either, since my ROCm build predates the architecture. I picked the UD-IQ3_XXS quant at 97.05 GiB over every 2-bit option because the experts ship natively at FP4, so 2-bit drops below native expert precision while 3-bit stays close to it. The 2-bit build would save about 7 GiB and both fit in 128 GB regardless. The full suite took 282 minutes.
- RADV dominates prompt processing across all model families (50-280% faster than AMDVLK on pp). Token generation is typically tied between drivers since it's bandwidth-bound. AMDVLK only wins on GPT-OSS-120B pp (661 vs 596 RADV).
- Kimi-Linear-48B was the early Polyglot leader at 22/65 before the server-based runs moved the board. Muse Glimmer now leads at 64, DeepSeek-V4-Flash follows at 63, Ornith reaches 57, Agents-A1 reaches 54, and Gemma-4-31B reaches 44. Kimi remains notable for a perfect Go rule engine and 72 tokens per second at 28 GiB.
- Devstral-Small-2 punches above its weight. A 24B dense coding model hitting 29/57 Postgres and 31/56 Cassandra while scoring 10/10 LRU and 59/59 LeetCode. The 15 t/s generation speed hurts but the quality-per-parameter is impressive.
- Nemotron-3-Nano collapses without thinking mode. 0/57 Postgres, 0/56 Cassandra, 0/59 LeetCode. The Mamba-2 hybrid architecture appears to need reasoning tokens enabled to produce structured output. Fast (1106 pp, 68 tg) but unusable for coding or database tasks with no-think.
- Magistral-Small reaches 35/56 on Cassandra, behind only DeepSeek-V4-Flash at 43 and the Gemma models among everything tested, and ahead of Devstral (31) and Kimi (24). But it scores 0/10 LRU, 12/57 Postgres, and leaked its reasoning scaffold into the creative-writing test output. A specialist, not a generalist.
- GLM-4.7-Flash can't code. 0/10 LRU, 0/8 FastAPI, 16/59 LeetCode (extraction failures), 0/65 polyglot. But it handles database work (23/57 PG, 27/56 Cass) and generates at 73 t/s. A fast model with a very narrow skill set.
- Laguna S 2.1 is Poolside's 118B mixture of experts with roughly 8B active parameters. I tested the official Q4_K_M on Poolside's llama.cpp fork because upstream support has not landed. Combined 198/285 puts it between Qwen3.6-35B-A3B and the Qwen3.5 models. The Combined run uses thinking for code and database work, then suppresses it for writing. It is much better at code and diagnosis than writing. LRU, FastAPI, and LeetCode are perfect, Polyglot reaches 35/65, and the hallucination set is a perfect 22/22. Creative writing scores 21/30 and the separate prose constraint test scores 53/80. RADV is the practical backend at 263 pp and 21.63 tg. ROCm reaches 305 pp with flash attention disabled but falls to 19.08 tg. Poolside's official DFlash draft is a firm reject on Strix Halo. It accepts only 10.7 percent of drafts on the fixed prompt set and cuts generation from 18.95 to 7.24 tokens per second.
- Agents-A1 moves the single-shot board but not the agentic one. The official Q8_0 reaches 1010 pp and 53.79 tg on RADV, then scores 224/285 Combined. Its 54/65 Polyglot score held the local record until the Ruby challenge was fixed, and it is now third. Agents-A1 is one of the few models whose Sinatra code fails on its own merits rather than because of that broken challenge. Tooling reaches 59.5/65 at 96.9 percent usable reliability and 7.1 seconds per call. In multi-turn Sokoban it stops at phase 4 in both modes after 53 minutes, 354 turns, and 342 tool calls. Writing remains mid-pack at 22/30. Its strongest roles are Polyglot and tooling, while Qwen3.6-27B remains the coding-agent pick.
- Puzzle 75B-A9B has a severe Strix thread cliff. ROCm reaches 22.66 tg at ten threads and only 2.16 at thirty-two. Combined reaches 200/285 with a strong 27/30 writing sample, 34/56 Cassandra, and a perfect 22/22 calibration run. The trade is weak 47.5/65 tooling, one successful prompt injection, a custom llama.cpp port, and one observed ROCm crash after a long server run. The complete benchmark record preserves the result.
- Ornith-1.0-35B is DeepReinforce's agentic-coding RL fine-tune of Qwen3.5-35B-A3B, and it is the strongest result that base has produced on this board. Combined 236/285 puts it third overall behind DeepSeek-V4-Flash and Gemma-4-31B, clearing its own base and the other agentic tune Nex-N2-mini by a wide margin. I ran it thinking-on, its native reasoning mode, the same way I ran Nex, so that gap reads cleanly as what the reinforcement training bought. The headline is polyglot at 57/65, second on the board. It got there by solving two challenges nothing local had cracked, the cron matcher at 10/10 and the sliding-window rate limiter at 9/10, with a perfect Go rule engine alongside. Its Ruby Sinatra webhook scores 9/10, not the zero I first recorded. I originally wrote that it produced real code with real logic bugs, and that was wrong. The challenge was broken. Cassandra is the soft spot at 28/56, almost all of it in Tier 4 procedural CQL, the same syntax failures the Nex tune had. It also returned a perfect 22/22 on the hallucination calibration set, matching the base rather than beating it, and a board-topping 77/80 on the prose constraint test where it held first-person present tense with no stray colons, semicolons, or dashes. I added a --think flag to the bench harness to drive its always-on reasoning properly, and ran the suite at a 32k token budget so the long reasoning trace had room to finish before the answer.
- MiniMax-M2.5 went from 0/10 coding to 10/10 after switching to Unsloth's UD-Q3_K_XL quant and their recommended sampling params (temp 1.0, top_p 0.95, min_p 0.01, top_k 40). At 94 GiB it is the largest artifact in this benchmark set, but 185/285 Combined puts it 6th overall. Perfect 10/10 on both PostgreSQL T2 optimization and Cassandra T2 anti-pattern detection. AMDVLK is its best backend at 32 t/s (vs 22 on RADV).
- MiniMax-M2.7 is the 229B successor with an UD-IQ4_XS quant that squeezes into 101 GiB. Writing improved to 28/30 and all coding benchmarks max out (10/10 LRU, 59/59 LeetCode, 8/8 FastAPI). The model has an unusual no-think quirk. Its M2.5 predecessor required a custom Jinja template to suppress thinking. M2.7 reverses this entirely. The standard --reasoning-budget 0 flag works, but the custom template causes the model to reason inline in plain text without ever producing code. Combined 179/285 puts it just behind M2.5 at 185. PostgreSQL T2 optimization and T3 diagnosis are both perfect (9/10 and 6/6). Cassandra T4 procedural challenges benefit from leaving thinking on (7/30 vs 2/30 without). RADV is the right backend here, unlike M2.5 which preferred AMDVLK.
- Qwen3.6-35B-A3B is the hybrid-attention successor to the 3.5 MoE. Combined 205/285, writing 29/30, and polyglot 35/65 put it near the top, but the 35 polyglot score is best-of-5 rather than stable single-run behavior. Database regressions come from Tier 4 procedural SQL and CQL, while raw writing output appends internal planning notes even with
enable_thinking:false. - Qwen3.6-27B is the dense companion to the 35B-A3B MoE. Combined 213/285 and writing 30/30 make it stronger on quality, while 12 tg makes it roughly five times slower than the MoE sibling. Polyglot averages 30/65 with Go variance, Postgres reaches 44/57, and the raw writing output has the same planning-note leak.
- Granite-4.1-30B ties MiniMax-M2.7 for the highest non-Qwen writing score at 28/30 and lands Combined 172/285. It also hits 10/10 LRU, 59/59 LeetCode, and perfect Postgres T2, but FastAPI stays at 2/8 and generation is slow at 11.76 tg. The 8B is notable for a perfect 22/22 hallucination-calibration run, while the 3B remains a speed-first model with LRU and FastAPI failures.
- Mistral-Medium-3.5-128B lands Combined 207/285 and writes 30/30, but coding needs the model's high-effort reasoning interface. Default no-reasoning LRU scored 2/10, while
reasoning_effort: highraised it to 7/10 with a different expiry-map bug. ROCm is the right backend for prompt processing, and token generation stays around 3 t/s across all backends because the 125B dense workload is memory-bandwidth-bound. - Nemotron-3-Nano-Omni-30B-A3B-Reasoning adds multimodal training but does not improve the text-only writing score. Combined 148/285 puts it below the working field, and task-specific sampling is required. Instruct-mode params raise LRU from 0/10 to 8/10 but lower polyglot from 6/65 to 2/65, while FastAPI and Cassandra T4 still fail on endpoint design and CQL syntax.
- Sampling params transform quality results. Unsloth-recommended params (presence_penalty 1.5, top_k 20, thinking mode via chat-template-kwargs) made the difference between 0/10 and 10/10 on several models.
- The Sinatra webhook is solved and was never really unsolved. It read as the one impossible Polyglot challenge for 104 runs because the harness was broken, and once that was fixed several models turned out to have been writing correct Ruby all along. Muse Glimmer scores it 10/10, which is independent confirmation that the fix holds on a model that never ran against the broken version. Agents-A1 remains one of the few that fails it on its own merits.
- Gemma 4 still leads both database benchmarks. Dense 31B tops the tracked PostgreSQL and Cassandra scores, while the 26B MoE follows closely. Agents-A1 reaches 40/57 on PostgreSQL and Puzzle reaches 34/56 on Cassandra.
- Mistral-Small-4-119B is the first Mistral mixture of experts on the board, 119B total with 6B active. At RADV it runs 363 pp and 40 tg, A6B-class generation for a 119B model and far quicker than the dense Mistral-Medium-3.5 at 3 tg. Combined 184/285 rests on strong pure coding (10/10 LRU, 59/59 LeetCode, 34/65 polyglot) against a mid-pack 21/30 on creative writing. Reasoning is binary, either none or high, and high is token-hungry. The first LRU run scored 0/10 because the reasoning trace filled the entire 4096-token budget before any code reached the response, so the suite runs at --max-tokens 20480. The creative-writing output ends with an assistant sign-off offering to revise the story, which the scorecard treats as a deduction.
- Nex-N2-mini is an agentic fine-tune of Qwen3.5-35B-A3B that tracks its base closely. Combined 183/285 sits 9 points under the 192 the base records. Writing, LRU, FastAPI, and polyglot match cell for cell. Postgres climbs 6 to 38/57, but Cassandra drops 9 to 29/56 on Tier 4 CQL syntax errors, and LeetCode loses the heap-merge problem. The agentic training only really shows in the one column the base never ran, automation reliability, where it scores 57/65 at 98.5% usable across the 65 calls with a perfect 100% on the JSON-envelope tier. I left thinking on for the whole suite because the server used the model's own template, so these numbers reflect its native reasoning mode rather than a suppressed one. The disk quant is s-batman's Q8_0 at 34.4 GiB, the best documented of the community GGUFs since no Unsloth build exists.
- The Cassandra JSON extractor now parses with strict=False, so anti-pattern answers that pretty-print multi-line CQL inside a JSON string value still count. Models often write a CREATE TABLE across several lines with literal newlines, which strict JSON rejects even when the answer is correct. The change raised Mistral-Small-4 Tier 2 from 5/10 to 7/10. Cassandra scores recorded before 2026-06-05 predate it and may understate any model that emitted the same multi-line JSON.
Partial Results
Models evaluated before the full benchmark suite was established. These ran writing and LRU cache tests but not the complete battery. Listed here for historical reference.
Performance
| Model | Quant | Size | RADV pp | RADV tg | AMDVLK pp | AMDVLK tg |
|---|---|---|---|---|---|---|
| Step3.5-Flash | IQ3_XS | 76 GiB | 237 | 32 | n/a | n/a |
| Nemotron-3-Super-120B-A12B | Unsloth UD-Q4_K_XL | 78 GiB | 196 | 10.2 | 139 | 9.86 |
Quality
| Model | Writing /30 | LRU /10 |
|---|---|---|
| Ling-Flash-2.0 | 26 | 2 |
| Nemotron-3-Super-120B-A12B | 25 | 10 |
| Devstral-2-123B | 25 | 2 |
| Solar-Open-100B | 21 | 0 |
| Mistral-Large-2411 | 20 | 2 |