← Back to Léo's Pages

Local models · dual DGX Sparks · TP2

How the household models actually score

Not a synthetic leaderboard. This is every Sparks build we have measured on Léo Laporte’s Hermes agent harness — real tools, mail and calendar cases, buried-needle reads, fairness under load, and decode speed — as of 16 September 2026.

Columns are model + quant + recipe. Rows are tests. A dash means we did not run that test on that build, never a fail. Gold cells are the best measured value in the row.

Quality

GLM still wins

The gap is precision traps (6/10 vs 2–4) and long-context needles (3/3 vs 0–1). Everything else is inside noise.

Speed

Qwen NVFP4 is the hare

It wins every throughput row and holds 12× the KV cache. Official FP8 gives that back for +1 corpus point.

Live seat

Qwen FP8, test drive

Resident since 16 Sep 09:09. GLM rung C is parked. Read Table A row 1, Table B needles, Table D decode together — that is the whole trade.

How to read this

Every number comes from the same dual NVIDIA DGX Spark cluster, tensor-parallel 2. The agentic corpus is 78 cases at effort low with a 4,096-token floor. Where a build ran more than once, the cell shows min–max across arms (run-to-run noise is about ±2 cases). “Carried” means the number is from an earlier run of the same weights on this recipe lineage.

GLM · Mia (HEAD → rung C)
GLM-5.3-Flash EXL3 4 bpw. Mia recipe lineage from 9 Sep through fair-v5 rung C (15 Sep). Parked 16 Sep.
GLM · Reederey
Same weights, Reederey87’s self-built ExLlamaV3 1.4.7 recipe. Bake-off 10 Sep, not adopted.
GLM · Mia c190db1
The 1–7 Sep resident (MTP k=2). Four corpus arms on an older harness.
DSV4.1 · EXL3
DeepSeek-V4.1-Flash EXL3 2.9 bpw, shipped recipe, text-only. 13 Sep.
FN · NVFP4
Qwen3.8-Flash-Next, NVIDIA NVFP4 checkpoint, Mia dual recipe. 13 Sep.
FN · FP8 resident
Qwen3.8-Flash-Next, Qwen’s official FP8 checkpoint. Resident since 16 Sep, test drive.

Table A — agentic corpus, 78 cases

Pass / cases, by suite. Bar ≥ 59 on the total. Visual strip uses the high end of any range.

78-case agentic corpus at effort low
Test (cases) GLM · Mia
HEAD → rung C
GLM · Reederey
2b5dfdc
GLM · Mia
c190db1
DSV4.1 · EXL3 FN · NVFP4 FN · FP8
TOTAL (78) — bar ≥ 59 60–62 59 57–61 55 55 56
gate-a1: agent basics, tool use, extraction (28) 21–22 22 19–22 19 21 21
gate-a2: general knowledge, refusal shape (10) 9 8 8–9 9 10 10
gate-a3: precision traps — unicode twin, dump bait, number decoys, negation, date (10) 6 6 6 2 3 4
calmail: JMAP / iCal / mail-rule cases (5) 2–4 2 3–4 4 3 3
coordination: multi-agent handoffs, hallucinated-execution trap (5) 3–4 3 3–4 3 4 3
maintenance: ops and housekeeping (4) 1–2 2 1–2 2 2 2
robustness: needles, determinism, concurrent prefill, canary (8) 8 8 7–8 8 4 5
twitads: TWiT-Ads domain cases (6) 6 6 5–6 6 6 6
webpage (2) 2 2 2 2 2 2

Qwen edges GLM on general knowledge (10/10) and once on coordination. The official FP8 checkpoint does not change Qwen’s shape: +1 total, a reshuffle of the same suites.

Table B — the cases that separate the models

Pass / fail, or k/n across arms. This is where GLM’s quality lead actually lives.

Separator cases
Case GLM · Mia rung C GLM · Reederey GLM · Mia c190db1 DSV4.1 · EXL3 FN · NVFP4 FN · FP8
rob-03 needle buried at 50K
rob-04 needle buried at 100K
rob-05 needle buried at 200K1/4 †
rob-06 determinism (same answer twice)
rob-02 concurrent prefill
a3-unicode-twin
a3-dump-bait
a3-number-decoys
a3-negation-trap
a3-date-operative
a1-todo-add-one
a1-extract-two-urls
coordination-02 hallucinated-execution trap2/33/4
calmail-03 JMAP false-success1/3
calmail-04 mail-rule recall over precision2/32/4

† c190db1 arms ran on harness b24e30c before the needle refit; the 200K needle there is not comparable.

Table C — other quality and reliability gates

Hammer, vision, soak, older ratings
Test GLM · Mia rung C GLM · Reederey GLM · Mia c190db1 DSV4.1 · EXL3 FN · NVFP4 FN · FP8 DSV4F · NVFP4
rob-01 hammer: 40 sequential thinking calls @4096 — content / empty / length 40/40 · 0 · 0 40/40 · 0 · 0 0 empty 40/40 · 0 · 0 40/40 · 0 · 1 length 40/40 · 0 · 0 0/40 empty @high
rob-01 p50 wall (lower is better) 24.1 s 30.2 s 58.2 s 29.7 s 45.0 s 26.7 s
Buried instruction survives 50K / 100K / 200K 3/3 3/3 2/3 † 3/3 0/3 1/3
Vision 12-image smoke (21 sub-cases since 13 Sep) pass (carried) none on GB10 21/21 21/21 never gated
Thinking-off probe (no reasoning leaks) clean leaked, later fixed clean clean
Léo’s 23-problem quality rating /100 (30 Aug) 89.5 ± 1.0 stock · 86.8 abliterated 91 NVFP4 GLM MTP-4, one trial
Coding eval /100 (29 Aug; frontier pair = 100) 80.4 on Sparks · 88.6 via z.ai cloud

Reference points on the same coding axis: frontier pair (Kronk / Pi) = 100; Mac mini qwen3-coder = 53.5. Effort vocabularies differ by build — a rejected medium or high is a 400, not a quality fail.

Table D — speed and capacity

Tests down, models across. Decode is tokens/s. Prefill is wall time to first content. Fairness bar is a short request waiting behind a long read, target < 10 s. Newcomer bar is first token of a 5.5K prompt ten seconds into another thinking reply, target < 20 s.

Throughput, prefill, fairness, KV
Test GLM rung C
15 Sep
GLM PR182
14 Sep
GLM 9348755
11–13 Sep
GLM HEAD
9 Sep
Reederey 2b5 Reederey 7d8 GLM c190db1 DSV4.1 EXL3 FN NVFP4 FN FP8 DSV4F NVFP4
Decode ×1 no-think (tok/s) 26.926.926.8 · 26.425.621.0 · 20.623.024.235.932.538–41
Decode ×1 thinking (tok/s) 34.034.032.6 · 31.034.233.7 ‡20.6 · 21.126.626.037.433.238–41
×4 aggregate (per stream) 62.7 (15.7)43.641.6 · 40.139.9 (10.0)52.543.9 · 42.537.732.2 (8.1)98.8 (24.7)75.5 (18.9)79.0
×8 aggregate (per stream) 59.5 (7.4)50.747.0 · 45.646.1 (5.8)52.144.1 · 42.945.332.1 (4.0)162.6 (20.3)111.5 (13.9)84.0
Prefill: 30K document to first content 38.7 s64.0 s19.4 s
Prefill: 80K / 128K 99.7 / 159 s173 / 282 s57.7 / 99 s
Fairness: short request behind a 30K read (bar < 10 s) 5.2 s4.9 s6.7 s65 s8.3 s10.6 s
Fairness: behind 80K / 128K 4.9 s / —5.2 s / —6.6 / 6.5 s181K: 202 s173 / 282 s8.4 / 7.8 s10.6 / 9.9 s
Newcomer: first token of a 5.5K prompt 10 s into another thinking reply (bar < 20 s) 12–14 s83–99 s
KV pool at 262K (tokens · concurrency) 643,072 · 2.45×643,072643,072 · 2.45×447,354 · 1.71×719,585 · 2.75×763,165 · 2.91×3,176,269 · 12×918,776 · 3.5×
Host memory headroom while serving ~5 GiB~5 GiB5 GiB5 GiB2–7 GiB10–12 GiB~3 GiB
Boot to /health 10.7 min9 min~13 min16.8 min
Status 16 Sep parkedparkedparkedsupersedednot adoptedsupersededsupersededparkedparkedRESIDENTcancelled 6 Sep

‡ Reederey 2b5dfdc single-stream thinking figure carried from the 10 Sep session summary; approximate. Only fair-v5 GLM (rung C) is under 5 s behind a long read and under 15 s for a newcomer during decode. Qwen FP8 sits right at the 10 s fairness line.

Table E — late-August GLM serving lineage

Older harness; corpus at high effort; rob-01 at 2048 and 4096. Kept so the September board has a floor.

August GLM recipes
Test EXL3 stock, kpool-fixed EXL3 abliterated EXL3 tip + MTP k=2 NVFP4 RedHat W4A4 + MTP-4 NVFP4 RedHat + DFlash2 NVFP4 ModelOpt + MTP-4 / DFlash2
Corpus 78, single-stream4950
Corpus 78, 4 requests at once4953
rob-01 empty @2048 (lower better)2320266–9 / 22
rob-01 empty @40960000
Decode ×1 (tok/s)23.624.619.5
×4 per stream / aggregate13.9–15.4 / 43.511.5–12.7 / 36.09.8–10.5 / 29.8
Léo’s quality /10089.5 ± 1.086.8 ± 1.091 (one trial)
Verdictresident 30 Aug → 1 Sepnot adoptedsupersedednot adoptednot adoptedhistorical

Table F — checked, but no comparable numbers

Out of board
ModelWhenWhat we knowStatus
Inkling NVFP4 (dual Sparks)20 AugFirst to pass all 13 evals of the early battery, then unusable in real work (hallucinated)dropped
Qwen3.8-Flash-Next NVFP4, single-Spark5 Seprob-01 21/40 and 23/40 blank; token-0 loop killed both containersdropped; 13 Sep dual superseded it
Qwen3.8-Flash FP8 via stock vLLM26 AugSmoke-loaded the day the weights dropped; no gatesuperseded 13 / 16 Sep
MiniMax H3AugRan on the Mac mini, not the Sparksout of scope
Qwen3.8-27B Uncensored Q4_K_M (Mojo, llama.cpp)ongoingVision, 128K; Hermes aux seat; never through the Sparks batteryfleet seat
Qwen3.6 MLX (Mac mini)ongoing64 GB unified; Hermes fallback seatfleet seat
nvidia/GLM-5.3-Flash-NVFP4 (official)watchModelOpt 0.47 W4A4 experts + dense MLP; not run herewatch list

Reading the board — 16 September 2026