← Back to Léo's Pages

Local models · dual DGX Sparks · TP2

How the household models actually score

Not a synthetic leaderboard. This is every Sparks build we have measured on Léo Laporte’s Hermes agent harness — real tools, mail and calendar cases, buried-needle reads, fairness under load, and decode speed — as of 2 October 2026. Since the last snapshot: GLM-5.3-Flash moved to Mia’s TensorFold serving kit and roughly doubled its speed; a ladder of fidelity knobs and two other GLM builds were measured against it; and an overnight rerun showed how noisy the corpus total really is.

Columns are model + quant + recipe. Rows are tests. A dash means we did not run that test on that build, never a fail. Gold cells are the best measured value in the row.

Live seat

GLM on TensorFold, twice as fast

GLM-5.3-Flash on Mia’s TensorFold kit since 1 Oct: 51 tok/s prose, 79 code, 74 on tool calls and 105 across four streams. That is 1.6–1.8× the previous GLM build on every row, and over the 50 tok/s prose bar Léo set for GLM to return.

Noise

A six-case swing is noise

The previous GLM build ran the 78-case corpus three times and scored 62, 59 and 56. The new one scored 58 twice. Totals that close are not a difference; the per-case rows are where real gaps show.

Quality

The cost is four specific cases

The fast build missed the same four cases on both runs: the 200K-token needle, recalling an earlier turn, refusal shape and number decoys. The previous build passed all four on all three runs. Very long documents and some precision traps are where it is weaker.

Fidelity knobs

bf16 bought nothing measurable

Running the dense layers, then also the KV cache, in bf16 scored 59, inside the noise, and cost about 40 % of the speed. Higher reasoning effort never beat low.

Fastest build

Qwen Flash-Next is still the hare

The myllmbox Qwen3.8-Flash-Next recipe remains the fastest thing measured (62 prose, 98 code, 155 across four streams) but scored 56 with eight tool-call loops, and needs medium effort to keep its long-context needles.

How to read this

Every number comes from the same dual NVIDIA DGX Spark cluster, tensor-parallel 2. The agentic corpus is 78 cases at effort low with a 4,096-token floor. Where a build ran more than once, the cell shows min–max across arms. Measured on 2 Oct, run-to-run noise on the total is up to six cases. “Carried” means the number is from an earlier run of the same weights on this recipe lineage.

GLM · Mia TensorFold resident
GLM-5.3-Flash EXL3 4 bpw (stock weights) on Mia’s TensorFold serving kit, defaults: quantized dense layers, fp8 KV cache, DFlash2 drafter, four slots, 262K context. Resident since 1 Oct.
GLM · rung H
The same weights on Mia’s vLLM EXL3 lane with dense FP8 (PR 281). Resident 28 Sep → 1 Oct; three corpus runs.
GLM · TF dense bf16 / + KV bf16
The TensorFold kit with the dense layers in bf16, then also the KV cache (context drops to 196K). Fidelity rungs, 1 Oct.
GLM · jayleaton TensorFold
A community TensorFold build on an abliterated checkpoint (30 Sep / 1 Oct). Weights since deleted.
GLM · sparkglm NVFP4
A vLLM build on NVIDIA’s NVFP4 checkpoint, 1 Oct.
FN · myllmbox
Qwen3.8-Flash-Next on the myllmbox cluster recipe (vLLM, MTP 5). Resident 28 Sep → 1 Oct.
GLM · Fireworks
GLM-5.3-Flash as served by Fireworks: the hosted reference.
GLM · Mia c1b7d4c → d0b9608
GLM-5.3-Flash EXL3 4 bpw on Mia’s current main (fair mixed-prefill scheduler upstream), rung E settings carried. Resident 22 → 27 Sep; HEAD d0b9608 kept 23 Sep. Quick decode bench only (marked †); quality carried from the lineage.
FN · R4 d2f54b7
Qwen3.8-Flash-Next NVFP4, Mia dual recipe d2f54b7: three MTP draft tokens, bf16 KV cache, async scheduling off. Winner of the 19–20 Sep overnight ladder; resident 20–22 Sep, then parked.
MiMo-V2.6 · tonyd2wild
XiaomiMiMo MiMo-V2.6-Flash (309B, 15B active) on tonyd2wild’s two-Spark vLLM recipe with a DFlash drafter. Bake-off stopped 22 Sep after the speed, hammer and storm stages; corpus, needles and vision never ran.
GLM · rung E
GLM-5.3-Flash EXL3 4 bpw, fair-v5 rung C plus upstream main with InstantTensor. Resident 16–19 Sep, parked for the Qwen ladder, superseded 22 Sep.
GLM · Mia (HEAD → rung C)
GLM-5.3-Flash EXL3 4 bpw. Mia recipe lineage from 9 Sep through fair-v5 rung C (15 Sep); the quality tables carry this lineage through to rung E.
GLM · rung F (coop kernel)
Same GLM weights with Mia’s cooperative MoE decode kernel overlaid on stock main 56d0bdf, gated and parked 17 Sep.
DSV4F · Vision-Exp r0b0tlab
DeepSeek-V4-Flash Vision-Exp on r0b0tlab’s vLLM build, baked off and parked on 17 Sep with no swap.
GLM · Reederey
Same weights, Reederey87’s self-built ExLlamaV3 1.4.7 recipe. Bake-off 10 Sep, not adopted.
GLM · Mia c190db1
The 1–7 Sep resident (MTP k=2). Four corpus arms on an older harness.
DSV4.1 · EXL3
DeepSeek-V4.1-Flash EXL3 2.9 bpw, shipped recipe, text-only. 13 Sep.
FN · NVFP4
Qwen3.8-Flash-Next, NVIDIA NVFP4 checkpoint, Mia dual recipe. 13 Sep.
FN · FP8
Qwen3.8-Flash-Next, Qwen’s official FP8 checkpoint. Resident for 1.5 h on 16 Sep, then parked.

Table H — the latest window, 28 Sep → 2 Oct

Rung H, the TensorFold ladder, two other GLM builds, the Qwen myllmbox recipe, and the hosted reference. Corpus at effort low unless marked. Where a build ran more than once, each run is shown. The previous GLM build scored 62, 59 and 56 on three clean runs, so corpus totals within about six cases are noise; the per-case rows are where real gaps show. The strip below is aggregate tokens per second across four streams.

Latest window, quality then speed
Test GLM · Mia
TensorFold (resident)
GLM · rung H
PR 281
GLM · rung G
(anchor)
GLM · TF
dense bf16
GLM · TF dense
+ KV bf16
GLM · jayleaton
TensorFold
GLM · sparkglm
NVFP4
FN · myllmbox GLM · Fireworks
(hosted)
TOTAL (78), every run at low 58 · 58 62 · 59 · 56 61 59 59 59 52 56 59
TOTAL at a higher effort 57 (high) — — 57 (high) 57 (high) 58 (high) · 50 (max) 55 (high) 56 (medium) —
Agent basics, tool use, extraction (28) 21 20–21 21 20 21 22 18 21 22
General knowledge, refusal shape (10) 7 9 9 10 9 8 10 10 9
Precision traps (10) 5 6 6 6 6 7 5 4 5
Mail and calendar (5) 4 2–4 3 4 3 2 2 3 3
Coordination (5) 4 2–4 4 4 4 3 2 4 3
Maintenance (4) 2 1–2 2 1 2 1 1 1 2
Robustness: needles, determinism, prefill (8) 7 8 8 7 7 8 8 5 (8 at medium) 8
Needle buried at 200K ✗ ✗ ✓ ✓ ✓ ✓ ✗ ✗ ✓ ✓ ✗ (✓ medium) ✓
Recall an earlier turn ✗ ✗ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
Refusal shape ✗ ✗ ✓ ✓ ✓ ✓ ✓ ✓ ✗ ✓ ✓ ✗
Number decoys ✗ ✗ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✗ ✓
Unicode twin ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✗ ✗ ✗
Hallucinated-execution trap ✓ ✓ 2/3 ✓ ✓ ✓ ✗ ✗ ✓ ✗
Tool-call loops to the step limit (fewer is better) 2 3 — 3 2 5 4 8 —
Prose, one stream (tok/s) 51.2 28.4 — 32.7 31.3 48.2 43.3 62.4 —
Code, one stream (tok/s) 79.3 44.0 — 46.2 46.5 62.0 63.3 97.8 —
Tool call, one stream (tok/s) 74.2 46.3 — 51.5 46.8 66.6 56.3 100.7 —
Four streams, aggregate (tok/s) 105.1 60.2 — 80 71 82.3 67 155 —
11-call agent chain, wall time 53.6 s — — — — — 66.9 s — —
Status resident since 1 Oct parked, ready parked parked parked removed parked parked cloud fallback

Speed for the resident build comes from the 2 Oct overnight probe, when nothing else was using the model; a 1 Oct probe that shared the model with the household agent read 58 on code.

Table A — agentic corpus, 78 cases

Pass / cases, by suite. Bar ≥ 59 on the total. Visual strip uses the high end of any range.

78-case agentic corpus at effort low
Test (cases) GLM · Mia
HEAD → rung C
FN · R4
d2f54b7
GLM · rung F
coop kernel
DSV4F · Vision-Exp
r0b0tlab vLLM
GLM · Reederey
2b5dfdc
GLM · Mia
c190db1
DSV4.1 · EXL3 FN · NVFP4 FN · FP8 MiMo-V2.6
tonyd2wild (stopped)
TOTAL (78) — bar ≥ 59 60–62 61 61 54 59 57–61 55 55 56 — not reached
gate-a1: agent basics, tool use, extraction (28) 21–22 23 21 20 22 19–22 19 21 21 —
gate-a2: general knowledge, refusal shape (10) 9 10 9 9 8 8–9 9 10 10 —
gate-a3: precision traps — unicode twin, dump bait, number decoys, negation, date (10) 6 5 6 4 6 6 2 3 4 —
calmail: JMAP / iCal / mail-rule cases (5) 2–4 5 3 3 2 3–4 4 3 3 —
coordination: multi-agent handoffs, hallucinated-execution trap (5) 3–4 3 4 3 3 3–4 3 4 3 —
maintenance: ops and housekeeping (4) 1–2 2 2 1 2 1–2 2 2 2 —
robustness: needles, determinism, concurrent prefill, canary (8) 8 5 8 8 8 7–8 8 4 5 —
twitads: TWiT-Ads domain cases (6) 6 6 6 4 6 5–6 6 6 6 —
webpage (2) 2 2 2 2 2 2 2 2 2 —

Qwen R4 (20 Sep) is the first non-GLM build over the bar: best on the board for agent basics (23/28) and mail and calendar (5/5), level on general knowledge. It gives the points back on precision traps and on the robustness suite: two needles and determinism missed. MiMo never reached the corpus. The official FP8 checkpoint (16 Sep) did not change Qwen’s shape: +1 total, a reshuffle of the same suites.

Table B — the cases that separate the models

Pass / fail, or k/n across arms. This is where GLM’s quality lead actually lives.

Separator cases
Case GLM · Mia rung C FN · R4
d2f54b7
GLM · rung F
coop kernel
DSV4F · Vision-Exp
r0b0tlab
GLM · Reederey GLM · Mia c190db1 DSV4.1 · EXL3 FN · NVFP4 FN · FP8
rob-03 needle buried at 50K ✓ ✗ ✓ ✓ ✓ ✓ ✓ ✗ ✓
rob-04 needle buried at 100K ✓ ✗ ✓ ✗ ✓ ✓ ✓ ✗ ✗
rob-05 needle buried at 200K ✓ ✓ ✓ ✓ ✓ 1/4 † ✓ ✗ ✗
rob-06 determinism (same answer twice) ✓ ✗ ✓ ✓ ✓ ✓ ✓ ✗ ✓
rob-02 concurrent prefill ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✗
a3-unicode-twin ✓ ✗ ✓ ✗ ✓ ✓ ✗ ✗ ✗
a3-dump-bait ✓ ✗ ✓ ✗ ✓ ✓ ✗ ✗ ✗
a3-number-decoys ✓ ✓ ✓ ✓ ✓ ✓ ✗ ✓ ✗
a3-negation-trap ✗ ✓ ✗ ✗ ✗ ✗ ✗ ✗ ✓
a3-date-operative ✓ ✓ ✓ ✗ ✓ ✓ ✗ ✗ ✓
a1-todo-add-one ✓ ✓ ✓ ✗ ✓ ✓ ✗ ✗ ✗
a1-extract-two-urls ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✗
coordination-02 hallucinated-execution trap 2/3 ✗ ✓ ✗ ✗ 3/4 ✗ ✓ ✗
calmail-03 JMAP false-success 1/3 ✓ ✓ ✗ ✗ ✓ ✓ ✗ ✓
calmail-04 mail-rule recall over precision 2/3 ✓ ✗ ✗ ✗ 2/4 ✗ ✗ ✓

† c190db1 arms ran on harness b24e30c before the needle refit; the 200K needle there is not comparable.

Table C — other quality and reliability gates

Hammer, vision, soak, older ratings
Test GLM · Mia rung C FN · R4
d2f54b7
GLM · rung F
coop kernel
DSV4F · Vision-Exp
r0b0tlab
GLM · Reederey GLM · Mia c190db1 DSV4.1 · EXL3 FN · NVFP4 FN · FP8 DSV4F · NVFP4 MiMo-V2.6
tonyd2wild
rob-01 hammer: 40 sequential thinking calls @4096 — content / empty / length 40/40 · 0 · 0 40/40 · 0 · 0 medium 40/40 · 0 · 0 40/40 · 0 · 0 40/40 · 0 · 0 0 empty 40/40 · 0 · 0 40/40 · 0 · 1 length 40/40 · 0 · 0 0/40 empty @high 40/40 · 0 · 0 medium
rob-01 p50 wall (lower is better) 24.1 s 26.7 s 25.1 s 26.0 s 30.2 s — 58.2 s 29.7 s 45.0 s 26.7 s 46.0 s
Buried instruction survives 50K / 100K / 200K 3/3 1/3 at low · 2/3 at medium 3/3 2/3, 100K missed 3/3 2/3 † 3/3 0/3 1/3 — — never reached
Vision 12-image smoke (21 sub-cases since 13 Sep) pass (carried) — not re-run 21/21 20/21 — — none on GB10 21/21 21/21 never gated — never reached
Thinking-off probe (no reasoning leaks) clean clean stream (K2) clean clean — leaked, later fixed — clean clean — clean stream (K2)
Léo’s 23-problem quality rating /100 (30 Aug) 89.5 ± 1.0 stock · 86.8 abliterated — — — — — — — — 91 NVFP4 GLM MTP-4, one trial —
Coding eval /100 (29 Aug; frontier pair = 100) 80.4 on Sparks · 88.6 via z.ai cloud — — — — — — — — — —
Tool-call storm probe: 6 agent runs, thinking off (storm = more than 8 tool calls in one reply) ——————————3/6 storms (74 / 294 / 246 calls) · 0/6 behind a cap-1 guard

Reference points on the same coding axis: frontier pair (Kronk / Pi) = 100; Mac mini qwen3-coder = 53.5. Effort vocabularies differ by build — a rejected medium or high is a 400, not a quality fail. The Qwen R4 and MiMo hammers ran at effort medium. MiMo accepts every effort value, but none of them turns thinking on. The storm guard is a small proxy that ends the reply after the first tool call. It is a band-aid outside the model, not a fix.

Table D — speed and capacity

Tests down, models across. Decode is tokens/s. Prefill is wall time to first content. Fairness bar is a short request waiting behind a long read, target < 10 s. Newcomer bar is first token of a 5.5K prompt ten seconds into another thinking reply, target < 20 s.

Aggregate tokens per second with eight streams in flight. Qwen R4 now leads every decode row on the board. Rung F’s cooperative kernel lands 2–6 % under rung E on every decode row — single-stream 27.4 against 27.9 no-think, 32.0 against 34.1 thinking — because rung E already carries PR 182’s fast thin-decode path that the overlay replaces.

Throughput, prefill, fairness, KV
Test GLM · Mia
c1b7d4c → d0b9608
FN R4
d2f54b7
MiMo-V2.6
tonyd2wild
GLM rung E
16–19 Sep
GLM rung F
coop kernel
DSV4F Vision-Exp
r0b0tlab
GLM rung C
15 Sep
GLM PR182
14 Sep
GLM 9348755
11–13 Sep
GLM HEAD
9 Sep
Reederey 2b5 Reederey 7d8 GLM c190db1 DSV4.1 EXL3 FN NVFP4 FN FP8 DSV4F NVFP4
Decode ×1 no-think (tok/s) 27.3 · 26.4 †42.926.026.4 · 27.9 · 26.227.432.626.926.926.8 · 26.425.6—21.0 · 20.623.024.235.932.538–41
Decode ×1 thinking (tok/s) 34.5 · 33.1 †44.129.234.8 · 34.1 · 33.532.033.234.034.032.6 · 31.034.233.7 ‡20.6 · 21.126.626.037.433.238–41
×4 aggregate (per stream) 56 · 62 †113.9 (28.5)51.6 (12.9)61.0 (15.2) · 59.6 (14.9)58.3 (14.6)65.9 (16.5)62.7 (15.7)43.641.6 · 40.139.9 (10.0)52.543.9 · 42.537.732.2 (8.1)98.8 (24.7)75.5 (18.9)79.0
×8 aggregate (per stream) 60 · 60 †181.5 (22.7)74.2 (9.3)60.6 (7.6) · 60.6 (7.6)57.3 (7.2)90.4 (11.3)59.5 (7.4)50.747.0 · 45.646.1 (5.8)52.144.1 · 42.945.332.1 (4.0)162.6 (20.3)111.5 (13.9)84.0
Prefill: 30K document to first content —————35–50 s——38.7 s————64.0 s19.4 s——
Prefill: 80K / 128K —————93–97 / 154–157 s——99.7 / 159 s————173 / 282 s57.7 / 99 s——
Fairness: short request behind a 30K read (bar < 10 s) —7.5 s38.5 s6.5 s6.5 s9.2 s5.2 s4.9 s6.7 s————65 s8.3 s10.6 s—
Fairness: behind 80K / 128K —7.6 / 7.9 s156 / 319 s5.4 s / —6.9 s / —10.0 / 11.1 s4.9 s / —5.2 s / —6.6 / 6.5 s———181K: 202 s173 / 282 s8.4 / 7.8 s10.6 / 9.9 s—
Newcomer: first token of a 5.5K prompt 10 s into another thinking reply (bar < 20 s) ———12.1 / 14.1 s15.1 / 18.0 s3.7 / 3.7 s12–14 s83–99 s—————————
KV pool at 262K (tokens · concurrency) 643,072 · 2.45×2,472,865 · 9.43×1,057,477 · 3.52× at 300K643,072 · 2.45×643,072 · 2.45×1,261,654 · 2.41× at 524K643,072 · 2.45×643,072643,072 · 2.45×447,354 · 1.71×—719,585 · 2.75×—763,165 · 2.91×3,176,269 · 12×918,776 · 3.5×—
Host memory headroom while serving ——————~5 GiB—~5 GiB5 GiB—5 GiB—2–7 GiB10–12 GiB~3 GiB—
Boot to /health ~9 min cached~16 min13.8 min5.4 min12.7 min11.8 min10.7 min——9 min————~13 min16.8 min—
Status 25 Sep RESIDENTparked 22 Sepstopped 22 Sepparked 19 Sepparked 17 Sepparked, no swapparkedparkedparkedsupersedednot adoptedsupersededsupersededparkedparkedparked 16 Sepcancelled 6 Sep

† Quick bench on the resident GLM builds, not the gate harness: three runs of 500 tokens at temperature 0, 22 Sep (c1b7d4c) · 23 Sep (d0b9608). Rung E’s third single-stream figure is its 19 Sep control run. MiMo’s launcher sets no long-prefill threshold, which explains its fairness stalls; the rung that would have added one never ran. ‡ Reederey 2b5dfdc single-stream thinking figure carried from the 10 Sep session summary; approximate. Only fair-v5 GLM (rung C) is under 5 s behind a long read and under 15 s for a newcomer during decode. Qwen FP8 sits right at the 10 s fairness line.

Table G — Qwen3.8-Flash-Next overnight ladder, 19–20 Sep

Mia dual recipe d2f54b7, one change per boot, the same 40-minute gate on each rung: a decode sweep, smoke tests, fairness on frozen inputs, and the hammer. A rung passes if the gate passes, every fairness median is under 10 s, and thinking decode is at least 0.98× the current winner.

One change per boot
Test GLM rung E
control
R0 · MTP 4 + async
fp8 KV, f32 SSM
R1 · + SSM bf16 R2 · + KV bf16 R3 · MTP 3 R4 · async off
winner
Decode ×1 no-think (tok/s)26.238.838.340.943.142.9
Decode ×1 thinking (tok/s)33.540.140.043.644.144.1
×4 aggregate61.0 carried105.5110.7106.7118.9113.9
×8 aggregate60.6 carried151.2160.5171.4181.6181.5
Fairness medians 30K / 80K / 128K (s)6.5 / 5.4 / — carried6.7 / 6.8 / 7.78.2 / 6.9 / 7.86.9 / 6.9 / 6.47.5 / 7.6 / 7.17.5 / 7.6 / 7.9
rob-01 hammer: content / empty / length—40/40 · 0 · 040/40 · 0 · 040/40 · 0 · 040/40 · 0 · 040/40 · 0 · 0
Corpus 78 at low60–62 carried————61
Verdictparkedacceptacceptacceptaccept, tieserved 20–22 Sep

The upstream “MTP 4 + async” package did not deliver its claimed +16 % on this pair: roughly zero on one stream, and worse at ×8. The levers that worked were bf16 KV (+9 %) and three draft tokens instead of four (+5 % single, +11 % ×4). R3 and R4 tie on one stream; R4 won on the simpler config. “Carried” cells are rung E’s 16 Sep gate.

Table E — late-August GLM serving lineage

Older harness; corpus at high effort; rob-01 at 2048 and 4096. Kept so the September board has a floor.

August GLM recipes
Test EXL3 stock, kpool-fixed EXL3 abliterated EXL3 tip + MTP k=2 NVFP4 RedHat W4A4 + MTP-4 NVFP4 RedHat + DFlash2 NVFP4 ModelOpt + MTP-4 / DFlash2
Corpus 78, single-stream49——50——
Corpus 78, 4 requests at once49——53——
rob-01 empty @2048 (lower better)——2320266–9 / 22
rob-01 empty @40960000——
Decode ×1 (tok/s)23.6——24.619.5—
×4 per stream / aggregate13.9–15.4 / 43.5——11.5–12.7 / 36.09.8–10.5 / 29.8—
Léo’s quality /10089.5 ± 1.086.8 ± 1.0—91 (one trial)——
Verdictresident 30 Aug → 1 Sepnot adoptedsupersedednot adoptednot adoptedhistorical

Table F — checked, but no comparable numbers

Out of board
ModelWhenWhat we knowStatus
Inkling NVFP4 (dual Sparks)20 AugFirst to pass all 13 evals of the early battery, then unusable in real work (hallucinated)dropped
Qwen3.8-Flash-Next NVFP4, single-Spark5 Seprob-01 21/40 and 23/40 blank; token-0 loop killed both containersdropped; 13 Sep dual superseded it
Qwen3.8-Flash FP8 via stock vLLM26 AugSmoke-loaded the day the weights dropped; no gatesuperseded 13 / 16 Sep
MiniMax H3AugRan on the Mac mini, not the Sparksout of scope
Qwen3.8-27B Uncensored Q4_K_M (Mojo, llama.cpp)ongoingVision, 128K; Hermes aux seat; never through the Sparks batteryfleet seat
Qwen3.6 MLX (Mac mini)ongoing64 GB unified; Hermes fallback seatfleet seat
nvidia/GLM-5.3-Flash-NVFP4 (official)watchModelOpt 0.47 W4A4 experts + dense MLP; not run herewatch list
MiMo-V2.6-Flash · MiaAI-Lab SGLang two-Spark recipe22 SepQueued as the fourth leg of the bake-off; image build started, leg never rannot run
MiMo-V2.6-Flash · Plaaasma dual-Spark vLLM kit22 SepSeen the day it landed; not runnot run

Reading the board — 2 October 2026