Quality
GLM still wins
The gap is precision traps (6/10 vs 2–4) and long-context needles (3/3 vs 0–1). Everything else is inside noise.
Local models · dual DGX Sparks · TP2
Not a synthetic leaderboard. This is every Sparks build we have measured on Léo Laporte’s Hermes agent harness — real tools, mail and calendar cases, buried-needle reads, fairness under load, and decode speed — as of 16 September 2026.
Quality
The gap is precision traps (6/10 vs 2–4) and long-context needles (3/3 vs 0–1). Everything else is inside noise.
Speed
It wins every throughput row and holds 12× the KV cache. Official FP8 gives that back for +1 corpus point.
Live seat
Resident since 16 Sep 09:09. GLM rung C is parked. Read Table A row 1, Table B needles, Table D decode together — that is the whole trade.
Every number comes from the same dual NVIDIA DGX Spark cluster, tensor-parallel 2. The agentic corpus is 78 cases at effort low with a 4,096-token floor. Where a build ran more than once, the cell shows min–max across arms (run-to-run noise is about ±2 cases). “Carried” means the number is from an earlier run of the same weights on this recipe lineage.
Pass / cases, by suite. Bar ≥ 59 on the total. Visual strip uses the high end of any range.
| Test (cases) | GLM · Mia HEAD → rung C |
GLM · Reederey 2b5dfdc |
GLM · Mia c190db1 |
DSV4.1 · EXL3 | FN · NVFP4 | FN · FP8 |
|---|---|---|---|---|---|---|
| TOTAL (78) — bar ≥ 59 | 60–62 | 59 | 57–61 | 55 | 55 | 56 |
| gate-a1: agent basics, tool use, extraction (28) | 21–22 | 22 | 19–22 | 19 | 21 | 21 |
| gate-a2: general knowledge, refusal shape (10) | 9 | 8 | 8–9 | 9 | 10 | 10 |
| gate-a3: precision traps — unicode twin, dump bait, number decoys, negation, date (10) | 6 | 6 | 6 | 2 | 3 | 4 |
| calmail: JMAP / iCal / mail-rule cases (5) | 2–4 | 2 | 3–4 | 4 | 3 | 3 |
| coordination: multi-agent handoffs, hallucinated-execution trap (5) | 3–4 | 3 | 3–4 | 3 | 4 | 3 |
| maintenance: ops and housekeeping (4) | 1–2 | 2 | 1–2 | 2 | 2 | 2 |
| robustness: needles, determinism, concurrent prefill, canary (8) | 8 | 8 | 7–8 | 8 | 4 | 5 |
| twitads: TWiT-Ads domain cases (6) | 6 | 6 | 5–6 | 6 | 6 | 6 |
| webpage (2) | 2 | 2 | 2 | 2 | 2 | 2 |
Qwen edges GLM on general knowledge (10/10) and once on coordination. The official FP8 checkpoint does not change Qwen’s shape: +1 total, a reshuffle of the same suites.
Pass / fail, or k/n across arms. This is where GLM’s quality lead actually lives.
| Case | GLM · Mia rung C | GLM · Reederey | GLM · Mia c190db1 | DSV4.1 · EXL3 | FN · NVFP4 | FN · FP8 |
|---|---|---|---|---|---|---|
| rob-03 needle buried at 50K | ✓ | ✓ | ✓ | ✓ | ✗ | ✓ |
| rob-04 needle buried at 100K | ✓ | ✓ | ✓ | ✓ | ✗ | ✗ |
| rob-05 needle buried at 200K | ✓ | ✓ | 1/4 † | ✓ | ✗ | ✗ |
| rob-06 determinism (same answer twice) | ✓ | ✓ | ✓ | ✓ | ✗ | ✓ |
| rob-02 concurrent prefill | ✓ | ✓ | ✓ | ✓ | ✓ | ✗ |
| a3-unicode-twin | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ |
| a3-dump-bait | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ |
| a3-number-decoys | ✓ | ✓ | ✓ | ✗ | ✓ | ✗ |
| a3-negation-trap | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ |
| a3-date-operative | ✓ | ✓ | ✓ | ✗ | ✗ | ✓ |
| a1-todo-add-one | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ |
| a1-extract-two-urls | ✓ | ✓ | ✓ | ✓ | ✓ | ✗ |
| coordination-02 hallucinated-execution trap | 2/3 | ✗ | 3/4 | ✗ | ✓ | ✗ |
| calmail-03 JMAP false-success | 1/3 | ✗ | ✓ | ✓ | ✗ | ✓ |
| calmail-04 mail-rule recall over precision | 2/3 | ✗ | 2/4 | ✗ | ✗ | ✓ |
† c190db1 arms ran on harness b24e30c before the needle refit; the 200K needle there is not comparable.
| Test | GLM · Mia rung C | GLM · Reederey | GLM · Mia c190db1 | DSV4.1 · EXL3 | FN · NVFP4 | FN · FP8 | DSV4F · NVFP4 |
|---|---|---|---|---|---|---|---|
| rob-01 hammer: 40 sequential thinking calls @4096 — content / empty / length | 40/40 · 0 · 0 | 40/40 · 0 · 0 | 0 empty | 40/40 · 0 · 0 | 40/40 · 0 · 1 length | 40/40 · 0 · 0 | 0/40 empty @high |
| rob-01 p50 wall (lower is better) | 24.1 s | 30.2 s | — | 58.2 s | 29.7 s | 45.0 s | 26.7 s |
| Buried instruction survives 50K / 100K / 200K | 3/3 | 3/3 | 2/3 † | 3/3 | 0/3 | 1/3 | — |
| Vision 12-image smoke (21 sub-cases since 13 Sep) | pass (carried) | — | — | none on GB10 | 21/21 | 21/21 | never gated |
| Thinking-off probe (no reasoning leaks) | clean | — | leaked, later fixed | — | clean | clean | — |
| Léo’s 23-problem quality rating /100 (30 Aug) | 89.5 ± 1.0 stock · 86.8 abliterated | — | — | — | — | — | 91 NVFP4 GLM MTP-4, one trial |
| Coding eval /100 (29 Aug; frontier pair = 100) | 80.4 on Sparks · 88.6 via z.ai cloud | — | — | — | — | — | — |
Reference points on the same coding axis: frontier pair (Kronk / Pi) = 100; Mac mini qwen3-coder = 53.5. Effort vocabularies differ by build — a rejected medium or high is a 400, not a quality fail.
Tests down, models across. Decode is tokens/s. Prefill is wall time to first content. Fairness bar is a short request waiting behind a long read, target < 10 s. Newcomer bar is first token of a 5.5K prompt ten seconds into another thinking reply, target < 20 s.
| Test | GLM rung C 15 Sep |
GLM PR182 14 Sep |
GLM 9348755 11–13 Sep |
GLM HEAD 9 Sep |
Reederey 2b5 | Reederey 7d8 | GLM c190db1 | DSV4.1 EXL3 | FN NVFP4 | FN FP8 | DSV4F NVFP4 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Decode ×1 no-think (tok/s) | 26.9 | 26.9 | 26.8 · 26.4 | 25.6 | — | 21.0 · 20.6 | 23.0 | 24.2 | 35.9 | 32.5 | 38–41 |
| Decode ×1 thinking (tok/s) | 34.0 | 34.0 | 32.6 · 31.0 | 34.2 | 33.7 ‡ | 20.6 · 21.1 | 26.6 | 26.0 | 37.4 | 33.2 | 38–41 |
| ×4 aggregate (per stream) | 62.7 (15.7) | 43.6 | 41.6 · 40.1 | 39.9 (10.0) | 52.5 | 43.9 · 42.5 | 37.7 | 32.2 (8.1) | 98.8 (24.7) | 75.5 (18.9) | 79.0 |
| ×8 aggregate (per stream) | 59.5 (7.4) | 50.7 | 47.0 · 45.6 | 46.1 (5.8) | 52.1 | 44.1 · 42.9 | 45.3 | 32.1 (4.0) | 162.6 (20.3) | 111.5 (13.9) | 84.0 |
| Prefill: 30K document to first content | — | — | 38.7 s | — | — | — | — | 64.0 s | 19.4 s | — | — |
| Prefill: 80K / 128K | — | — | 99.7 / 159 s | — | — | — | — | 173 / 282 s | 57.7 / 99 s | — | — |
| Fairness: short request behind a 30K read (bar < 10 s) | 5.2 s | 4.9 s | 6.7 s | — | — | — | — | 65 s | 8.3 s | 10.6 s | — |
| Fairness: behind 80K / 128K | 4.9 s / — | 5.2 s / — | 6.6 / 6.5 s | — | — | — | 181K: 202 s | 173 / 282 s | 8.4 / 7.8 s | 10.6 / 9.9 s | — |
| Newcomer: first token of a 5.5K prompt 10 s into another thinking reply (bar < 20 s) | 12–14 s | 83–99 s | — | — | — | — | — | — | — | — | — |
| KV pool at 262K (tokens · concurrency) | 643,072 · 2.45× | 643,072 | 643,072 · 2.45× | 447,354 · 1.71× | — | 719,585 · 2.75× | — | 763,165 · 2.91× | 3,176,269 · 12× | 918,776 · 3.5× | — |
| Host memory headroom while serving | ~5 GiB | — | ~5 GiB | 5 GiB | — | 5 GiB | — | 2–7 GiB | 10–12 GiB | ~3 GiB | — |
| Boot to /health | 10.7 min | — | — | 9 min | — | — | — | — | ~13 min | 16.8 min | — |
| Status 16 Sep | parked | parked | parked | superseded | not adopted | superseded | superseded | parked | parked | RESIDENT | cancelled 6 Sep |
‡ Reederey 2b5dfdc single-stream thinking figure carried from the 10 Sep session summary; approximate. Only fair-v5 GLM (rung C) is under 5 s behind a long read and under 15 s for a newcomer during decode. Qwen FP8 sits right at the 10 s fairness line.
Older harness; corpus at high effort; rob-01 at 2048 and 4096. Kept so the September board has a floor.
| Test | EXL3 stock, kpool-fixed | EXL3 abliterated | EXL3 tip + MTP k=2 | NVFP4 RedHat W4A4 + MTP-4 | NVFP4 RedHat + DFlash2 | NVFP4 ModelOpt + MTP-4 / DFlash2 |
|---|---|---|---|---|---|---|
| Corpus 78, single-stream | 49 | — | — | 50 | — | — |
| Corpus 78, 4 requests at once | 49 | — | — | 53 | — | — |
| rob-01 empty @2048 (lower better) | — | — | 23 | 20 | 26 | 6–9 / 22 |
| rob-01 empty @4096 | 0 | 0 | 0 | 0 | — | — |
| Decode ×1 (tok/s) | 23.6 | — | — | 24.6 | 19.5 | — |
| ×4 per stream / aggregate | 13.9–15.4 / 43.5 | — | — | 11.5–12.7 / 36.0 | 9.8–10.5 / 29.8 | — |
| Léo’s quality /100 | 89.5 ± 1.0 | 86.8 ± 1.0 | — | 91 (one trial) | — | — |
| Verdict | resident 30 Aug → 1 Sep | not adopted | superseded | not adopted | not adopted | historical |
| Model | When | What we know | Status |
|---|---|---|---|
| Inkling NVFP4 (dual Sparks) | 20 Aug | First to pass all 13 evals of the early battery, then unusable in real work (hallucinated) | dropped |
| Qwen3.8-Flash-Next NVFP4, single-Spark | 5 Sep | rob-01 21/40 and 23/40 blank; token-0 loop killed both containers | dropped; 13 Sep dual superseded it |
| Qwen3.8-Flash FP8 via stock vLLM | 26 Aug | Smoke-loaded the day the weights dropped; no gate | superseded 13 / 16 Sep |
| MiniMax H3 | Aug | Ran on the Mac mini, not the Sparks | out of scope |
| Qwen3.8-27B Uncensored Q4_K_M (Mojo, llama.cpp) | ongoing | Vision, 128K; Hermes aux seat; never through the Sparks battery | fleet seat |
| Qwen3.6 MLX (Mac mini) | ongoing | 64 GB unified; Hermes fallback seat | fleet seat |
| nvidia/GLM-5.3-Flash-NVFP4 (official) | watch | ModelOpt 0.47 W4A4 experts + dense MLP; not run here | watch list |