Not a synthetic leaderboard. This is every Sparks build we have measured on Léo Laporte’s Hermes agent harness — real tools, mail and calendar cases, buried-needle reads, fairness under load, and decode speed — as of 2 October 2026. Since the last snapshot: GLM-5.3-Flash moved to Mia’s TensorFold serving kit and roughly doubled its speed; a ladder of fidelity knobs and two other GLM builds were measured against it; and an overnight rerun showed how noisy the corpus total really is.
Columns are model + quant + recipe. Rows are tests. A dash means we did not run that test on that build, never a fail. Gold cells are the best measured value in the row.
Live seat
GLM on TensorFold, twice as fast
GLM-5.3-Flash on Mia’s TensorFold kit since 1 Oct: 51 tok/s prose, 79 code, 74 on tool calls and 105 across four streams. That is 1.6–1.8× the previous GLM build on every row, and over the 50 tok/s prose bar Léo set for GLM to return.
Noise
A six-case swing is noise
The previous GLM build ran the 78-case corpus three times and scored 62, 59 and 56. The new one scored 58 twice. Totals that close are not a difference; the per-case rows are where real gaps show.
Quality
The cost is four specific cases
The fast build missed the same four cases on both runs: the 200K-token needle, recalling an earlier turn, refusal shape and number decoys. The previous build passed all four on all three runs. Very long documents and some precision traps are where it is weaker.
Fidelity knobs
bf16 bought nothing measurable
Running the dense layers, then also the KV cache, in bf16 scored 59, inside the noise, and cost about 40 % of the speed. Higher reasoning effort never beat low.
Fastest build
Qwen Flash-Next is still the hare
The myllmbox Qwen3.8-Flash-Next recipe remains the fastest thing measured (62 prose, 98 code, 155 across four streams) but scored 56 with eight tool-call loops, and needs medium effort to keep its long-context needles.
How to read this
Every number comes from the same dual NVIDIA DGX Spark cluster, tensor-parallel 2. The agentic corpus is 78 cases at effort low with a 4,096-token floor. Where a build ran more than once, the cell shows min–max across arms. Measured on 2 Oct, run-to-run noise on the total is up to six cases. “Carried” means the number is from an earlier run of the same weights on this recipe lineage.
GLM · Mia TensorFoldresident
GLM-5.3-Flash EXL3 4 bpw (stock weights) on Mia’s TensorFold serving kit, defaults: quantized dense layers, fp8 KV cache, DFlash2 drafter, four slots, 262K context. Resident since 1 Oct.
GLM · rung H
The same weights on Mia’s vLLM EXL3 lane with dense FP8 (PR 281). Resident 28 Sep → 1 Oct; three corpus runs.
GLM · TF dense bf16 / + KV bf16
The TensorFold kit with the dense layers in bf16, then also the KV cache (context drops to 196K). Fidelity rungs, 1 Oct.
GLM · jayleaton TensorFold
A community TensorFold build on an abliterated checkpoint (30 Sep / 1 Oct). Weights since deleted.
GLM · sparkglm NVFP4
A vLLM build on NVIDIA’s NVFP4 checkpoint, 1 Oct.
FN · myllmbox
Qwen3.8-Flash-Next on the myllmbox cluster recipe (vLLM, MTP 5). Resident 28 Sep → 1 Oct.
GLM · Fireworks
GLM-5.3-Flash as served by Fireworks: the hosted reference.
GLM · Mia c1b7d4c → d0b9608
GLM-5.3-Flash EXL3 4 bpw on Mia’s current main (fair mixed-prefill scheduler upstream), rung E settings carried. Resident 22 → 27 Sep; HEAD d0b9608 kept 23 Sep. Quick decode bench only (marked †); quality carried from the lineage.
FN · R4 d2f54b7
Qwen3.8-Flash-Next NVFP4, Mia dual recipe d2f54b7: three MTP draft tokens, bf16 KV cache, async scheduling off. Winner of the 19–20 Sep overnight ladder; resident 20–22 Sep, then parked.
MiMo-V2.6 · tonyd2wild
XiaomiMiMo MiMo-V2.6-Flash (309B, 15B active) on tonyd2wild’s two-Spark vLLM recipe with a DFlash drafter. Bake-off stopped 22 Sep after the speed, hammer and storm stages; corpus, needles and vision never ran.
GLM · rung E
GLM-5.3-Flash EXL3 4 bpw, fair-v5 rung C plus upstream main with InstantTensor. Resident 16–19 Sep, parked for the Qwen ladder, superseded 22 Sep.
GLM · Mia (HEAD → rung C)
GLM-5.3-Flash EXL3 4 bpw. Mia recipe lineage from 9 Sep through fair-v5 rung C (15 Sep); the quality tables carry this lineage through to rung E.
GLM · rung F (coop kernel)
Same GLM weights with Mia’s cooperative MoE decode kernel overlaid on stock main 56d0bdf, gated and parked 17 Sep.
DSV4F · Vision-Exp r0b0tlab
DeepSeek-V4-Flash Vision-Exp on r0b0tlab’s vLLM build, baked off and parked on 17 Sep with no swap.
GLM · Reederey
Same weights, Reederey87’s self-built ExLlamaV3 1.4.7 recipe. Bake-off 10 Sep, not adopted.
GLM · Mia c190db1
The 1–7 Sep resident (MTP k=2). Four corpus arms on an older harness.
Qwen3.8-Flash-Next, NVIDIA NVFP4 checkpoint, Mia dual recipe. 13 Sep.
FN · FP8
Qwen3.8-Flash-Next, Qwen’s official FP8 checkpoint. Resident for 1.5 h on 16 Sep, then parked.
Table H — the latest window, 28 Sep → 2 Oct
Rung H, the TensorFold ladder, two other GLM builds, the Qwen myllmbox recipe, and the hosted reference. Corpus at effort low unless marked. Where a build ran more than once, each run is shown. The previous GLM build scored 62, 59 and 56 on three clean runs, so corpus totals within about six cases are noise; the per-case rows are where real gaps show. The strip below is aggregate tokens per second across four streams.
FN · myllmbox
155
GLM · Mia TensorFold
105.1
GLM · jayleaton TF
82.3
GLM · TF dense bf16
80
GLM · TF dense + KV bf16
71
GLM · sparkglm
67
GLM · rung H
60.2
Latest window, quality then speed
Test
GLM · Mia TensorFold (resident)
GLM · rung H PR 281
GLM · rung G (anchor)
GLM · TF dense bf16
GLM · TF dense + KV bf16
GLM · jayleaton TensorFold
GLM · sparkglm NVFP4
FN · myllmbox
GLM · Fireworks (hosted)
TOTAL (78), every run at low
58 · 58
62 · 59 · 56
61
59
59
59
52
56
59
TOTAL at a higher effort
57 (high)
—
—
57 (high)
57 (high)
58 (high) · 50 (max)
55 (high)
56 (medium)
—
Agent basics, tool use, extraction (28)
21
20–21
21
20
21
22
18
21
22
General knowledge, refusal shape (10)
7
9
9
10
9
8
10
10
9
Precision traps (10)
5
6
6
6
6
7
5
4
5
Mail and calendar (5)
4
2–4
3
4
3
2
2
3
3
Coordination (5)
4
2–4
4
4
4
3
2
4
3
Maintenance (4)
2
1–2
2
1
2
1
1
1
2
Robustness: needles, determinism, prefill (8)
7
8
8
7
7
8
8
5 (8 at medium)
8
Needle buried at 200K
✗ ✗
✓ ✓ ✓
✓
✗
✗
✓
✓
✗ (✓ medium)
✓
Recall an earlier turn
✗ ✗
✓ ✓ ✓
✓
✓
✓
✓
✓
✓
✓
Refusal shape
✗ ✗
✓ ✓ ✓
✓
✓
✓
✗
✓
✓
✗
Number decoys
✗ ✗
✓ ✓ ✓
✓
✓
✓
✓
✓
✗
✓
Unicode twin
✓ ✓
✓ ✓ ✓
✓
✓
✓
✓
✗
✗
✗
Hallucinated-execution trap
✓ ✓
2/3
✓
✓
✓
✗
✗
✓
✗
Tool-call loops to the step limit (fewer is better)
2
3
—
3
2
5
4
8
—
Prose, one stream (tok/s)
51.2
28.4
—
32.7
31.3
48.2
43.3
62.4
—
Code, one stream (tok/s)
79.3
44.0
—
46.2
46.5
62.0
63.3
97.8
—
Tool call, one stream (tok/s)
74.2
46.3
—
51.5
46.8
66.6
56.3
100.7
—
Four streams, aggregate (tok/s)
105.1
60.2
—
80
71
82.3
67
155
—
11-call agent chain, wall time
53.6 s
—
—
—
—
—
66.9 s
—
—
Status
resident since 1 Oct
parked, ready
parked
parked
parked
removed
parked
parked
cloud fallback
Speed for the resident build comes from the 2 Oct overnight probe, when nothing else was using the model; a 1 Oct probe that shared the model with the household agent read 58 on code.
Table A — agentic corpus, 78 cases
Pass / cases, by suite. Bar ≥ 59 on the total. Visual strip uses the high end of any range.
GLM · Mia rung C
60–62
FN · R4
61
GLM · rung F
61
GLM · Mia c190db1
57–61
GLM · Reederey
59
FN · FP8
56
FN · NVFP4
55
DSV4.1 · EXL3
55
DSV4F · r0b0tlab
54
78-case agentic corpus at effort low
Test (cases)
GLM · Mia HEAD → rung C
FN · R4 d2f54b7
GLM · rung F coop kernel
DSV4F · Vision-Exp r0b0tlab vLLM
GLM · Reederey 2b5dfdc
GLM · Mia c190db1
DSV4.1 · EXL3
FN · NVFP4
FN · FP8
MiMo-V2.6 tonyd2wild (stopped)
TOTAL (78) — bar ≥ 59
60–62
61
61
54
59
57–61
55
55
56
— not reached
gate-a1: agent basics, tool use, extraction (28)
21–22
23
21
20
22
19–22
19
21
21
—
gate-a2: general knowledge, refusal shape (10)
9
10
9
9
8
8–9
9
10
10
—
gate-a3: precision traps — unicode twin, dump bait, number decoys, negation, date (10)
Qwen R4 (20 Sep) is the first non-GLM build over the bar: best on the board for agent basics (23/28) and mail and calendar (5/5), level on general knowledge. It gives the points back on precision traps and on the robustness suite: two needles and determinism missed. MiMo never reached the corpus. The official FP8 checkpoint (16 Sep) did not change Qwen’s shape: +1 total, a reshuffle of the same suites.
Table B — the cases that separate the models
Pass / fail, or k/n across arms. This is where GLM’s quality lead actually lives.
Separator cases
Case
GLM · Mia rung C
FN · R4 d2f54b7
GLM · rung F coop kernel
DSV4F · Vision-Exp r0b0tlab
GLM · Reederey
GLM · Mia c190db1
DSV4.1 · EXL3
FN · NVFP4
FN · FP8
rob-03 needle buried at 50K
✓
✗
✓
✓
✓
✓
✓
✗
✓
rob-04 needle buried at 100K
✓
✗
✓
✗
✓
✓
✓
✗
✗
rob-05 needle buried at 200K
✓
✓
✓
✓
✓
1/4 †
✓
✗
✗
rob-06 determinism (same answer twice)
✓
✗
✓
✓
✓
✓
✓
✗
✓
rob-02 concurrent prefill
✓
✓
✓
✓
✓
✓
✓
✓
✗
a3-unicode-twin
✓
✗
✓
✗
✓
✓
✗
✗
✗
a3-dump-bait
✓
✗
✓
✗
✓
✓
✗
✗
✗
a3-number-decoys
✓
✓
✓
✓
✓
✓
✗
✓
✗
a3-negation-trap
✗
✓
✗
✗
✗
✗
✗
✗
✓
a3-date-operative
✓
✓
✓
✗
✓
✓
✗
✗
✓
a1-todo-add-one
✓
✓
✓
✗
✓
✓
✗
✗
✗
a1-extract-two-urls
✓
✓
✓
✓
✓
✓
✓
✓
✗
coordination-02 hallucinated-execution trap
2/3
✗
✓
✗
✗
3/4
✗
✓
✗
calmail-03 JMAP false-success
1/3
✓
✓
✗
✗
✓
✓
✗
✓
calmail-04 mail-rule recall over precision
2/3
✓
✗
✗
✗
2/4
✗
✗
✓
† c190db1 arms ran on harness b24e30c before the needle refit; the 200K needle there is not comparable.
Reference points on the same coding axis: frontier pair (Kronk / Pi) = 100; Mac mini qwen3-coder = 53.5. Effort vocabularies differ by build — a rejected medium or high is a 400, not a quality fail. The Qwen R4 and MiMo hammers ran at effort medium. MiMo accepts every effort value, but none of them turns thinking on. The storm guard is a small proxy that ends the reply after the first tool call. It is a band-aid outside the model, not a fix.
Table D — speed and capacity
Tests down, models across. Decode is tokens/s. Prefill is wall time to first content. Fairness bar is a short request waiting behind a long read, target < 10 s. Newcomer bar is first token of a 5.5K prompt ten seconds into another thinking reply, target < 20 s.
FN · R4
181.5
FN · NVFP4
162.6
FN · FP8
111.5
DSV4F · Vision-Exp
90.4
MiMo-V2.6 · tonyd2wild
74.2
GLM · rung E
60.6
GLM · Mia d0b9608 †
60
GLM · rung F coop kernel
57.3
DSV4.1 · EXL3
32.1
Aggregate tokens per second with eight streams in flight. Qwen R4 now leads every decode row on the board. Rung F’s cooperative kernel lands 2–6 % under rung E on every decode row — single-stream 27.4 against 27.9 no-think, 32.0 against 34.1 thinking — because rung E already carries PR 182’s fast thin-decode path that the overlay replaces.
Throughput, prefill, fairness, KV
Test
GLM · Mia c1b7d4c → d0b9608
FN R4 d2f54b7
MiMo-V2.6 tonyd2wild
GLM rung E 16–19 Sep
GLM rung F coop kernel
DSV4F Vision-Exp r0b0tlab
GLM rung C 15 Sep
GLM PR182 14 Sep
GLM 9348755 11–13 Sep
GLM HEAD 9 Sep
Reederey 2b5
Reederey 7d8
GLM c190db1
DSV4.1 EXL3
FN NVFP4
FN FP8
DSV4F NVFP4
Decode ×1 no-think (tok/s)
27.3 · 26.4 †
42.9
26.0
26.4 · 27.9 · 26.2
27.4
32.6
26.9
26.9
26.8 · 26.4
25.6
—
21.0 · 20.6
23.0
24.2
35.9
32.5
38–41
Decode ×1 thinking (tok/s)
34.5 · 33.1 †
44.1
29.2
34.8 · 34.1 · 33.5
32.0
33.2
34.0
34.0
32.6 · 31.0
34.2
33.7 ‡
20.6 · 21.1
26.6
26.0
37.4
33.2
38–41
×4 aggregate (per stream)
56 · 62 †
113.9 (28.5)
51.6 (12.9)
61.0 (15.2) · 59.6 (14.9)
58.3 (14.6)
65.9 (16.5)
62.7 (15.7)
43.6
41.6 · 40.1
39.9 (10.0)
52.5
43.9 · 42.5
37.7
32.2 (8.1)
98.8 (24.7)
75.5 (18.9)
79.0
×8 aggregate (per stream)
60 · 60 †
181.5 (22.7)
74.2 (9.3)
60.6 (7.6) · 60.6 (7.6)
57.3 (7.2)
90.4 (11.3)
59.5 (7.4)
50.7
47.0 · 45.6
46.1 (5.8)
52.1
44.1 · 42.9
45.3
32.1 (4.0)
162.6 (20.3)
111.5 (13.9)
84.0
Prefill: 30K document to first content
—
—
—
—
—
35–50 s
—
—
38.7 s
—
—
—
—
64.0 s
19.4 s
—
—
Prefill: 80K / 128K
—
—
—
—
—
93–97 / 154–157 s
—
—
99.7 / 159 s
—
—
—
—
173 / 282 s
57.7 / 99 s
—
—
Fairness: short request behind a 30K read (bar < 10 s)
—
7.5 s
38.5 s
6.5 s
6.5 s
9.2 s
5.2 s
4.9 s
6.7 s
—
—
—
—
65 s
8.3 s
10.6 s
—
Fairness: behind 80K / 128K
—
7.6 / 7.9 s
156 / 319 s
5.4 s / —
6.9 s / —
10.0 / 11.1 s
4.9 s / —
5.2 s / —
6.6 / 6.5 s
—
—
—
181K: 202 s
173 / 282 s
8.4 / 7.8 s
10.6 / 9.9 s
—
Newcomer: first token of a 5.5K prompt 10 s into another thinking reply (bar < 20 s)
—
—
—
12.1 / 14.1 s
15.1 / 18.0 s
3.7 / 3.7 s
12–14 s
83–99 s
—
—
—
—
—
—
—
—
—
KV pool at 262K (tokens · concurrency)
643,072 · 2.45×
2,472,865 · 9.43×
1,057,477 · 3.52× at 300K
643,072 · 2.45×
643,072 · 2.45×
1,261,654 · 2.41× at 524K
643,072 · 2.45×
643,072
643,072 · 2.45×
447,354 · 1.71×
—
719,585 · 2.75×
—
763,165 · 2.91×
3,176,269 · 12×
918,776 · 3.5×
—
Host memory headroom while serving
—
—
—
—
—
—
~5 GiB
—
~5 GiB
5 GiB
—
5 GiB
—
2–7 GiB
10–12 GiB
~3 GiB
—
Boot to /health
~9 min cached
~16 min
13.8 min
5.4 min
12.7 min
11.8 min
10.7 min
—
—
9 min
—
—
—
—
~13 min
16.8 min
—
Status 25 Sep
RESIDENT
parked 22 Sep
stopped 22 Sep
parked 19 Sep
parked 17 Sep
parked, no swap
parked
parked
parked
superseded
not adopted
superseded
superseded
parked
parked
parked 16 Sep
cancelled 6 Sep
† Quick bench on the resident GLM builds, not the gate harness: three runs of 500 tokens at temperature 0, 22 Sep (c1b7d4c) · 23 Sep (d0b9608). Rung E’s third single-stream figure is its 19 Sep control run. MiMo’s launcher sets no long-prefill threshold, which explains its fairness stalls; the rung that would have added one never ran. ‡ Reederey 2b5dfdc single-stream thinking figure carried from the 10 Sep session summary; approximate. Only fair-v5 GLM (rung C) is under 5 s behind a long read and under 15 s for a newcomer during decode. Qwen FP8 sits right at the 10 s fairness line.
Table G — Qwen3.8-Flash-Next overnight ladder, 19–20 Sep
Mia dual recipe d2f54b7, one change per boot, the same 40-minute gate on each rung: a decode sweep, smoke tests, fairness on frozen inputs, and the hammer. A rung passes if the gate passes, every fairness median is under 10 s, and thinking decode is at least 0.98× the current winner.
One change per boot
Test
GLM rung E control
R0 · MTP 4 + async fp8 KV, f32 SSM
R1 · + SSM bf16
R2 · + KV bf16
R3 · MTP 3
R4 · async off winner
Decode ×1 no-think (tok/s)
26.2
38.8
38.3
40.9
43.1
42.9
Decode ×1 thinking (tok/s)
33.5
40.1
40.0
43.6
44.1
44.1
×4 aggregate
61.0 carried
105.5
110.7
106.7
118.9
113.9
×8 aggregate
60.6 carried
151.2
160.5
171.4
181.6
181.5
Fairness medians 30K / 80K / 128K (s)
6.5 / 5.4 / — carried
6.7 / 6.8 / 7.7
8.2 / 6.9 / 7.8
6.9 / 6.9 / 6.4
7.5 / 7.6 / 7.1
7.5 / 7.6 / 7.9
rob-01 hammer: content / empty / length
—
40/40 · 0 · 0
40/40 · 0 · 0
40/40 · 0 · 0
40/40 · 0 · 0
40/40 · 0 · 0
Corpus 78 at low
60–62 carried
—
—
—
—
61
Verdict
parked
accept
accept
accept
accept, tie
served 20–22 Sep
The upstream “MTP 4 + async” package did not deliver its claimed +16 % on this pair: roughly zero on one stream, and worse at ×8. The levers that worked were bf16 KV (+9 %) and three draft tokens instead of four (+5 % single, +11 % ×4). R3 and R4 tie on one stream; R4 won on the simpler config. “Carried” cells are rung E’s 16 Sep gate.
Table E — late-August GLM serving lineage
Older harness; corpus at high effort; rob-01 at 2048 and 4096. Kept so the September board has a floor.
August GLM recipes
Test
EXL3 stock, kpool-fixed
EXL3 abliterated
EXL3 tip + MTP k=2
NVFP4 RedHat W4A4 + MTP-4
NVFP4 RedHat + DFlash2
NVFP4 ModelOpt + MTP-4 / DFlash2
Corpus 78, single-stream
49
—
—
50
—
—
Corpus 78, 4 requests at once
49
—
—
53
—
—
rob-01 empty @2048 (lower better)
—
—
23
20
26
6–9 / 22
rob-01 empty @4096
0
0
0
0
—
—
Decode ×1 (tok/s)
23.6
—
—
24.6
19.5
—
×4 per stream / aggregate
13.9–15.4 / 43.5
—
—
11.5–12.7 / 36.0
9.8–10.5 / 29.8
—
Léo’s quality /100
89.5 ± 1.0
86.8 ± 1.0
—
91 (one trial)
—
—
Verdict
resident 30 Aug → 1 Sep
not adopted
superseded
not adopted
not adopted
historical
Table F — checked, but no comparable numbers
Out of board
Model
When
What we know
Status
Inkling NVFP4 (dual Sparks)
20 Aug
First to pass all 13 evals of the early battery, then unusable in real work (hallucinated)
dropped
Qwen3.8-Flash-Next NVFP4, single-Spark
5 Sep
rob-01 21/40 and 23/40 blank; token-0 loop killed both containers
dropped; 13 Sep dual superseded it
Qwen3.8-Flash FP8 via stock vLLM
26 Aug
Smoke-loaded the day the weights dropped; no gate
superseded 13 / 16 Sep
MiniMax H3
Aug
Ran on the Mac mini, not the Sparks
out of scope
Qwen3.8-27B Uncensored Q4_K_M (Mojo, llama.cpp)
ongoing
Vision, 128K; Hermes aux seat; never through the Sparks battery
fleet seat
Qwen3.6 MLX (Mac mini)
ongoing
64 GB unified; Hermes fallback seat
fleet seat
nvidia/GLM-5.3-Flash-NVFP4 (official)
watch
ModelOpt 0.47 W4A4 experts + dense MLP; not run here
Queued as the fourth leg of the bake-off; image build started, leg never ran
not run
MiMo-V2.6-Flash · Plaaasma dual-Spark vLLM kit
22 Sep
Seen the day it landed; not run
not run
Reading the board — 2 October 2026
The order is intelligence first, then reliability, then speed, with one refinement from late September: a small measured quality gap can lose to a large speed win, and GLM needed more than 50 tok/s on prose to return as the main model.
Resident: GLM-5.3-Flash on Mia’s TensorFold kit (since 1 Oct). 51 prose, 79 code, 74 tool calls, 105 across four streams, with the fewest tool-call loops on the board and an 11-call agent chain in 53.6 s.
The corpus total is noisier than we thought. Rung H scored 62, 59 and 56 on three clean runs. The TensorFold build’s 58 (twice) sits inside that band, so the total alone does not separate them.
The real difference is four cases, and it is consistent. The TensorFold build missed the 200K needle, recall of an earlier turn, refusal shape and number decoys on both runs; rung H passed all four every time. The bf16 dense rung gets the three short cases back but still misses the needle, which points at the kit for the needle and at its quantized dense layers for the rest (one run each, not proven).
Fidelity knobs: bf16 dense and bf16 dense + KV both scored 59, inside the noise, for about 40 % less speed. Higher effort never beat low.
Hosted reference: Fireworks’ GLM scored 59, inside the same band, so the local quants have not lost measurable intelligence; only speed differs.
Earlier on this board (25 Sep): Qwen R4 matched GLM’s total but lost the precision and needle rows and was replaced after two days in the seat; MiMo-V2.6-Flash was stopped after tool-call storms; tables A–G below keep those builds.