gezel Gezel Handboek

Model scorecard

These are measured results, not estimates. Each number is the count of jobs a model actually finished correctly, checked by a program on a real machine.

How we test models explains what was counted and what these numbers do not tell you. The short version: each job is run three times, a job passes only when every requirement is met, and runs lost to machine trouble are set aside rather than blamed on the model.

Results by test round

Each round has one table for each task set it ran, all measured on the same machine and build. Most rounds ran two sets. General capability covers writing a small working program, fixing a bug from its symptoms, following a procedure and stopping at a problem, and turning several documents into one reconciled summary. Office and knowledge work covers a customer notice under a hard word limit, a meeting turned into an action register, a cited research brief, an experiment read-out, a spreadsheet model, a slide deck, and a Word document. Some rounds add the two harder sets: Engineering work, which is developer work such as reviewing a change and proving a fix with tests, and Complex workflows, which asks a model to write and repair reusable recipes of its own.

Latest round — 2026-09-18

DGX Spark Class · NVIDIA GB10 · 122 GB · ds4 + llama-cpp · 3 trials per task · 2026-09-18 · gezel bcc9d87c · catalog 0.1.67

Memory isn't shown for this round. This machine's graphics chip shares the computer's memory in a way our measurement can't see, so the number would read far below what the model really uses.

General capability
ModelSizeTasks passedPerformanceContext
qwen3.8-flash-next-q2180B32/33 (97%)27.1 tok/s output
433 tok/s prefill
—
qwen3.8-flash-next-iq4 (kv: q4_0)180B31/33 (94%)25.8 tok/s output
480 tok/s prefill
64K

Not published — some tasks could not be measured on this round: qwen3.8-27b-q4 (1 task(s) unmeasured)

Earlier round — 2026-09-08

DGX Spark Class · NVIDIA GB10 · 122 GB · llama-cpp · 3 trials per task · 2026-09-08 · gezel 35c47b9f · catalog 0.1.56 — different gezel build, different catalog version

Memory isn't shown for this round. This machine's graphics chip shares the computer's memory in a way our measurement can't see, so the number would read far below what the model really uses.

General capability
ModelSizeTasks passedQualityPerformanceContext
qwen3.8-flash-next-iq4 (kv: q4_0)180B33/33 (100%)8.2/10 (9 pieces)24.5 tok/s output
380 tok/s prefill
64K
gemma4-31b-q4 (kv: f16)30.7B32/33 (97%)6.6/10 (9 pieces)35.9 tok/s output
484 tok/s prefill
90K
qwen3.8-27b-q4 (kv: f16)27B32/33 (97%)8.1/10 (8 pieces)28.6 tok/s output
752 tok/s prefill
256K
qwen3.6-35b-a3b-q4 (kv: f16)35B31/33 (94%)7.8/10 (9 pieces)87.4 tok/s output
1,675 tok/s prefill
64K
muse-glimmer-30b-q4 (kv: f16)30B27/33 (82%)6.8/10 (9 pieces)12.4 tok/s output
932 tok/s prefill
128K
nemotron3.5-lightning-30b-q4 (kv: f16)30B22/33 (67%)5.1/10 (8 pieces)103.2 tok/s output
1,705 tok/s prefill
64K
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContext
qwen3.6-35b-a3b-q4 (kv: f16)35B36/39 (92%)7.5/10 (26 pieces)87.4 tok/s output
1,675 tok/s prefill
64K
qwen3.8-27b-q4 (kv: f16)27B36/39 (92%)7.7/10 (26 pieces)28.6 tok/s output
752 tok/s prefill
256K
gemma4-31b-q4 (kv: f16)30.7B35/39 (90%)6.9/10 (25 pieces)35.9 tok/s output
484 tok/s prefill
90K
qwen3.8-flash-next-iq4 (kv: q4_0)180B35/39 (90%)8/10 (27 pieces)24.5 tok/s output
380 tok/s prefill
64K
muse-glimmer-30b-q4 (kv: f16)30B28/39 (72%)7.6/10 (26 pieces)12.4 tok/s output
932 tok/s prefill
128K
nemotron3.5-lightning-30b-q4 (kv: f16)30B19/39 (49%)6.3/10 (26 pieces)103.2 tok/s output
1,705 tok/s prefill
64K
Engineering work
ModelSizeTasks passedQualityPerformanceContext
gemma4-31b-q4 (kv: f16)30.7B20/30 (67%)5.1/10 (15 pieces)35.9 tok/s output
484 tok/s prefill
90K
qwen3.6-35b-a3b-q4 (kv: f16)35B17/27 (63%)6.7/10 (16 pieces)87.4 tok/s output
1,675 tok/s prefill
64K
qwen3.8-flash-next-iq4 (kv: q4_0)180B15/26 (58%)7.7/10 (14 pieces)24.3 tok/s output
420 tok/s prefill
64K
qwen3.8-27b-q4 (kv: f16)27B16/29 (55%)7.6/10 (16 pieces)28 tok/s output
749 tok/s prefill
256K
muse-glimmer-30b-q4 (kv: f16)30B13/30 (43%)5.7/10 (12 pieces)12.3 tok/s output
930 tok/s prefill
128K
nemotron3.5-lightning-30b-q4 (kv: f16)30B13/30 (43%)5/10 (10 pieces)103.2 tok/s output
1,705 tok/s prefill
64K
Complex workflows
ModelSizeTasks passedQualityPerformanceContext
gemma4-31b-q4 (kv: f16)30.7B17/27 (63%)2.7/10 (6 pieces)35.9 tok/s output
484 tok/s prefill
90K
qwen3.8-flash-next-iq4 (kv: q4_0)180B16/27 (59%)0.5/10 (6 pieces)24.3 tok/s output
416 tok/s prefill
64K
qwen3.6-35b-a3b-q4 (kv: f16)35B15/27 (56%)3/10 (6 pieces)87.4 tok/s output
1,675 tok/s prefill
64K
qwen3.8-27b-q4 (kv: f16)27B14/27 (52%)2.2/10 (6 pieces)28.5 tok/s output
753 tok/s prefill
256K
muse-glimmer-30b-q4 (kv: f16)30B13/27 (48%)0.5/10 (6 pieces)12.3 tok/s output
921 tok/s prefill
128K
nemotron3.5-lightning-30b-q4 (kv: f16)30B7/27 (26%)1.9/10 (6 pieces)103.2 tok/s output
1,705 tok/s prefill
64K

Earlier round — 2026-09-04

AMD Ryzen 9 7950X3D 16-Core Processor · AMD Radeon AI PRO R9700 · 64 GB · llama-cpp · 3 trials per task · 2026-09-04 · gezel f115c698 · catalog 0.1.55 — different device, different gezel build, different catalog version

System memory counts this computer's own memory only. What the model keeps on the separate graphics card isn't included.

General capability
ModelSizeTasks passedQualityPerformanceContextSystem memory
qwen3.5-27b-q4 (kv: f16)27B30/33 (91%)6.4/10 (9 pieces)55.9 tok/s output
656 tok/s prefill
159K22.2 GB
qwen3.5-9b-q4 (kv: f16)9B26/33 (79%)5.3/10 (9 pieces)122.9 tok/s output
2,143 tok/s prefill
256K8.9 GB
qwen3.5-4b-q4 (kv: f16)4B16/33 (48%)5.3/10 (9 pieces)153 tok/s output
3,349 tok/s prefill
256K7.2 GB
qwen3.5-2b-q4 (kv: f16)2B6/33 (18%)2.7/10 (9 pieces)203.3 tok/s output
7,605 tok/s prefill
256K5.1 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextSystem memory
qwen3.5-27b-q4 (kv: f16)27B36/39 (92%)6.9/10 (27 pieces)55.9 tok/s output
656 tok/s prefill
159K21.1 GB
qwen3.5-9b-q4 (kv: f16)9B32/39 (82%)6/10 (27 pieces)122.9 tok/s output
2,143 tok/s prefill
256K7.6 GB
qwen3.5-4b-q4 (kv: f16)4B28/39 (72%)5.3/10 (26 pieces)157 tok/s output
3,429 tok/s prefill
256K6.0 GB
qwen3.5-2b-q4 (kv: f16)2B20/39 (51%)4.3/10 (27 pieces)203.3 tok/s output
7,605 tok/s prefill
256K3.7 GB

Earlier round — 2026-09-03

M2 · Apple M2 (integrated GPU) · 16 GB · llama-cpp · 1 trial per task · 2026-09-03 · gezel a06b1070 · catalog 0.1.52 — different device, different gezel build, different catalog version, different trial count

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
gemma4-12b-q4 (kv: f16)12B10/11 (some tasks run once — count not rate)6.6/10 (3 pieces)13.7 tok/s output
100 tok/s prefill
147K10.1 GB
ornith1.5-9b-q4 (kv: f16)9B10/11 (some tasks run once — count not rate)6.5/10 (1 pieces)9.8 tok/s output
154 tok/s prefill
132K10.1 GB
qwen3.5-9b-q4 (kv: f16)9B9/11 (some tasks run once — count not rate)3.8/10 (1 pieces)9.2 tok/s output
121 tok/s prefill
126K10.8 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
ornith1.5-9b-q4 (kv: f16)9B12/13 (some tasks run once — count not rate)5.8/10 (9 pieces)9.8 tok/s output
154 tok/s prefill
132K10.0 GB
qwen3.5-9b-q4 (kv: f16)9B11/13 (some tasks run once — count not rate)5.3/10 (9 pieces)9.2 tok/s output
121 tok/s prefill
126K10.8 GB
gemma4-12b-q4 (kv: f16)12B10/13 (some tasks run once — count not rate)5/10 (9 pieces)13.7 tok/s output
100 tok/s prefill
147K10.0 GB

Earlier round — 2026-09-02

AMD Ryzen AI 9 HX 370 w/ Radeon 890M · NVIDIA GeForce RTX 5070 Ti Laptop GPU · 31 GB · llama-cpp · 1 trial per task · 2026-09-02 · gezel 7694351b · catalog 0.1.52 — different device, different gezel build, different catalog version, different trial count

System memory counts this computer's own memory only. What the model keeps on the separate graphics card isn't included.

General capability
ModelSizeTasks passedQualityPerformanceContextSystem memory
gemma4-12b-q4 (kv: f16)12B10/11 (some tasks run once — count not rate)7.4/10 (3 pieces)24.6 tok/s output
1,495 tok/s prefill
154K11.2 GB
qwen3.6-35b-a3b-q4 (kv: f16)35B10/11 (some tasks run once — count not rate)8.1/10 (3 pieces)15.1 tok/s output
80 tok/s prefill
64K21.0 GB
ornith1.5-9b-q4 (kv: f16)9B9/11 (some tasks run once — count not rate)7.2/10 (3 pieces)57.7 tok/s output
2,230 tok/s prefill
136K7.9 GB
qwen3.5-9b-q4 (kv: f16)9B9/11 (some tasks run once — count not rate)5.2/10 (3 pieces)29.9 tok/s output
329 tok/s prefill
130K9.8 GB
qwen3.8-27b-q2 (kv: f16)27B8/11 (some tasks run once — count not rate)7.9/10 (3 pieces)7 tok/s output
55 tok/s prefill
64K14.9 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextSystem memory
qwen3.5-9b-q4 (kv: f16)9B12/13 (some tasks run once — count not rate)6/10 (9 pieces)29.9 tok/s output
329 tok/s prefill
130K8.3 GB
qwen3.6-35b-a3b-q4 (kv: f16)35B12/13 (some tasks run once — count not rate)6.9/10 (9 pieces)15.1 tok/s output
80 tok/s prefill
64K21.1 GB
gemma4-12b-q4 (kv: f16)12B11/13 (some tasks run once — count not rate)5.2/10 (9 pieces)24.6 tok/s output
1,495 tok/s prefill
154K10.9 GB
ornith1.5-9b-q4 (kv: f16)9B11/13 (some tasks run once — count not rate)5.8/10 (9 pieces)57.7 tok/s output
2,230 tok/s prefill
136K7.5 GB
qwen3.8-27b-q2 (kv: f16)27B10/13 (some tasks run once — count not rate)6.8/10 (9 pieces)7 tok/s output
55 tok/s prefill
64K14.5 GB

Earlier round — 2026-08-31

DGX Spark Class · NVIDIA GB10 · 122 GB · ds4 + llama-cpp · 3 trials per task · 2026-08-31 · gezel f59825ea · catalog 0.1.51 — different gezel build, different catalog version

Memory isn't shown for this round. This machine's graphics chip shares the computer's memory in a way our measurement can't see, so the number would read far below what the model really uses.

General capability
ModelSizeTasks passedQualityPerformanceContext
deepseek-v4-flash-284b-q2284B33/33 (100%)8/10 (7 pieces)17.1 tok/s output
456 tok/s prefill
—
qwen3.8-27b-q4 (kv: f16)27B33/33 (100%)7.8/10 (9 pieces)31.1 tok/s output
708 tok/s prefill
256K
ornith1.5-35b-a3b-q8 (kv: f16)35B29/33 (88%)7.7/10 (7 pieces)42.8 tok/s output
1,248 tok/s prefill
64K
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContext
deepseek-v4-flash-284b-q2284B36/39 (92%)6.6/10 (27 pieces)17.1 tok/s output
456 tok/s prefill
—
ornith1.5-35b-a3b-q8 (kv: f16)35B36/39 (92%)7/10 (27 pieces)42.8 tok/s output
1,248 tok/s prefill
64K
qwen3.8-27b-q4 (kv: f16)27B36/39 (92%)7.2/10 (27 pieces)31.1 tok/s output
708 tok/s prefill
256K
Engineering work
ModelSizeTasks passedQualityPerformanceContext
ornith1.5-35b-a3b-q8 (kv: f16)35B19/30 (63%)6.1/10 (17 pieces)42.8 tok/s output
1,248 tok/s prefill
64K
qwen3.8-27b-q4 (kv: f16)27B18/30 (60%)6.9/10 (15 pieces)31.1 tok/s output
708 tok/s prefill
256K

Not published — some tasks could not be measured on this round: deepseek-v4-flash-284b-q2 (2 task(s) unmeasured)

Complex workflows
ModelSizeTasks passedQualityPerformanceContext
qwen3.8-27b-q4 (kv: f16)27B15/27 (56%)2.4/10 (6 pieces)29.1 tok/s output
713 tok/s prefill
256K
ornith1.5-35b-a3b-q8 (kv: f16)35B12/26 (46%)2.4/10 (6 pieces)42.8 tok/s output
1,248 tok/s prefill
64K

Not published — some tasks could not be measured on this round: deepseek-v4-flash-284b-q2 (1 task(s) unmeasured)

Earlier round — 2026-08-30

DGX Spark Class · NVIDIA GB10 · 122 GB · llama-cpp · 3 trials per task · 2026-08-30 · gezel 8d7bb7e5 · catalog 0.1.49 — different gezel build, different catalog version

Memory isn't shown for this round. This machine's graphics chip shares the computer's memory in a way our measurement can't see, so the number would read far below what the model really uses.

General capability
ModelSizeTasks passedQualityPerformanceContext
qwen3.8-27b-q4 (kv: f16)27B32/33 (97%)8.2/10 (9 pieces)32.4 tok/s output
777 tok/s prefill
256K
qwen3.6-35b-a3b-q4 (kv: f16)35B31/33 (94%)8/10 (9 pieces)82 tok/s output
1,690 tok/s prefill
64K
ornith1.5-9b-q8 (kv: f16)9B28/32 (88%)7.1/10 (9 pieces)25.6 tok/s output
2,117 tok/s prefill
256K
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContext
qwen3.8-27b-q4 (kv: f16)27B35/39 (90%)7.2/10 (27 pieces)32.4 tok/s output
777 tok/s prefill
256K
qwen3.6-35b-a3b-q4 (kv: f16)35B34/39 (87%)6.9/10 (26 pieces)82 tok/s output
1,690 tok/s prefill
64K
ornith1.5-9b-q8 (kv: f16)9B32/39 (82%)6.4/10 (27 pieces)25.6 tok/s output
2,117 tok/s prefill
256K

Earlier round — 2026-08-27

AMD Ryzen 9 7950X3D 16-Core Processor · AMD Radeon AI PRO R9700 · 64 GB · llama-cpp · 3 trials per task · 2026-08-27 · gezel 3cf0b7a2 · catalog 0.1.44 — different device, different gezel build, different catalog version

System memory counts this computer's own memory only. What the model keeps on the separate graphics card isn't included.

General capability
ModelSizeTasks passedQualityPerformanceContextSystem memory
qwen3.8-27b-q4 (kv: q8_0)27B32/33 (97%)8.1/10 (18 pieces)31.8 tok/s output
681 tok/s prefill
256K15.8 GB
qwen3.6-35b-a3b-q4 (kv: q8_0)35B30/33 (91%)7.8/10 (9 pieces)131.1 tok/s output
2,102 tok/s prefill
64K16.1 GB
gemma4-31b-q4 (kv: f16)30.7B29/33 (88%)7.1/10 (7 pieces)28.7 tok/s output
407 tok/s prefill
99K16.8 GB
gemma4-26b-q4 (kv: f16)25.2B28/33 (85%)6.9/10 (6 pieces)131.2 tok/s output
1,820 tok/s prefill
64K14.9 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextSystem memory
qwen3.8-27b-q4 (kv: q8_0)27B36/39 (92%)7.5/10 (26 pieces)31.8 tok/s output
681 tok/s prefill
256K14.9 GB
qwen3.6-35b-a3b-q4 (kv: q8_0)35B35/39 (90%)6.8/10 (26 pieces)131.1 tok/s output
2,102 tok/s prefill
64K15.6 GB
gemma4-31b-q4 (kv: f16)30.7B30/39 (77%)6.5/10 (20 pieces)28.7 tok/s output
407 tok/s prefill
99K16.9 GB
gemma4-26b-q4 (kv: f16)25.2B26/39 (67%)5.7/10 (22 pieces)131.2 tok/s output
1,820 tok/s prefill
64K14.7 GB

Earlier round — 2026-08-26

M4 Max · 64 GB · llama-cpp + mlx · 3 trials per task · 2026-08-26 · gezel d1d5e5c0 · catalog 0.1.42 — different device, different gezel build, different catalog version

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q2 (kv: q8_0)27B32/33 (97%)7.9/10 (9 pieces)24.1 tok/s output
241 tok/s prefill
256K21.1 GB
qwen3.8-27b-q3 (kv: q8_0)27B32/33 (97%)8/10 (9 pieces)22.7 tok/s output
239 tok/s prefill
256K24.5 GB
qwen3.8-27b-q4 (kv: q8_0)27B31/33 (94%)7.8/10 (17 pieces)21.9 tok/s output
232 tok/s prefill
256K28.4 GB
qwen3.8-27b-q427B31/33 (94%)7.8/10 (17 pieces)68.4 tok/s output
232 tok/s prefill
——
qwen3.8-27b-iq1-s (kv: q8_0)27B14/33 (42%)—11.1 tok/s output
190 tok/s prefill
256K21.1 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q4 (kv: q8_0)27B36/39 (92%)7.3/10 (53 pieces)21.9 tok/s output
232 tok/s prefill
256K26.2 GB
qwen3.8-27b-q427B36/39 (92%)7.3/10 (53 pieces)68.4 tok/s output
232 tok/s prefill
——
qwen3.8-27b-q2 (kv: q8_0)27B35/39 (90%)7.3/10 (25 pieces)24.1 tok/s output
241 tok/s prefill
256K19.8 GB
qwen3.8-27b-q3 (kv: q8_0)27B34/39 (87%)7.1/10 (27 pieces)22.7 tok/s output
239 tok/s prefill
256K24.4 GB
qwen3.8-27b-iq1-s (kv: q8_0)27B14/39 (36%)3.9/10 (25 pieces)11.1 tok/s output
190 tok/s prefill
256K20.2 GB

Earlier round — 2026-08-22

M4 Max · 64 GB · llama-cpp + mlx · 3 trials per task · 2026-08-22 · gezel b3be0a7f · catalog 0.1.39 — different device, different gezel build, different catalog version

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
gemma4-12b-q4 (kv: f16)12B31/33 (94%)6.9/10 (6 pieces)48.2 tok/s output
508 tok/s prefill
99K42.2 GB
gemma4-12b-q8 (kv: f16)12B30/33 (91%)8.1/10 (3 pieces)30.3 tok/s output
506 tok/s prefill
83K42.0 GB
btl4-compact-iq2 (kv: q8_0)35B29/33 (88%)7.1/10 (9 pieces)80.4 tok/s output
1,120 tok/s prefill
256K14.6 GB
ornith1.5-9b-q8 (kv: q8_0)9B28/33 (85%)4.4/10 (5 pieces)47 tok/s output
823 tok/s prefill
256K15.9 GB
qwen3.5-9b-q4 (kv: q8_0)9B28/33 (85%)5.3/10 (9 pieces)62.9 tok/s output
787 tok/s prefill
256K13.1 GB
gemma4-e4b-q4 (kv: f16)8B26/33 (79%)5.4/10 (9 pieces)96.5 tok/s output
1,442 tok/s prefill
128K13.7 GB
qwen3.5-4b-q4 (kv: q8_0)4B22/33 (67%)6.6/10 (9 pieces)89.3 tok/s output
1,336 tok/s prefill
256K11.5 GB
gemma4-e2b-q4 (kv: f16)2.3B17/33 (52%)4.3/10 (9 pieces)144.1 tok/s output
2,574 tok/s prefill
128K6.9 GB
lfm2.5-2.6b-q4 (kv: q8_0)2.6B13/33 (39%)3.6/10 (8 pieces)168.4 tok/s output
2,125 tok/s prefill
125K4.9 GB
qwen3.5-2b-q4 (kv: q8_0)2B9/33 (27%)3.2/10 (9 pieces)155.3 tok/s output
3,084 tok/s prefill
256K5.2 GB
mistral-7b-q4 (kv: q8_0)7B4/33 (12%)—76 tok/s output
762 tok/s prefill
32K9.6 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
gemma4-12b-q8 (kv: f16)12B35/39 (90%)5.8/10 (24 pieces)30.3 tok/s output
506 tok/s prefill
83K41.8 GB
gemma4-12b-q4 (kv: f16)12B34/39 (87%)5.8/10 (24 pieces)48.2 tok/s output
508 tok/s prefill
99K41.5 GB
ornith1.5-9b-q8 (kv: q8_0)9B34/39 (87%)6.5/10 (27 pieces)47 tok/s output
823 tok/s prefill
256K14.9 GB
gemma4-e4b-q4 (kv: f16)8B31/39 (79%)5.5/10 (22 pieces)96.5 tok/s output
1,442 tok/s prefill
128K12.5 GB
qwen3.5-9b-q4 (kv: q8_0)9B31/39 (79%)6.2/10 (25 pieces)62.9 tok/s output
787 tok/s prefill
256K11.3 GB
qwen3.5-4b-q4 (kv: q8_0)4B30/39 (77%)6.2/10 (24 pieces)89.3 tok/s output
1,336 tok/s prefill
256K9.9 GB
btl4-compact-iq2 (kv: q8_0)35B27/39 (69%)6.2/10 (25 pieces)80.4 tok/s output
1,120 tok/s prefill
256K14.1 GB
gemma4-e2b-q4 (kv: f16)2.3B16/39 (41%)4.5/10 (20 pieces)144.1 tok/s output
2,574 tok/s prefill
128K6.2 GB
lfm2.5-2.6b-q4 (kv: q8_0)2.6B13/39 (33%)4.4/10 (27 pieces)168.4 tok/s output
2,125 tok/s prefill
125K4.8 GB
qwen3.5-2b-q4 (kv: q8_0)2B10/39 (26%)4.1/10 (26 pieces)155.3 tok/s output
3,084 tok/s prefill
256K4.6 GB
mistral-7b-q4 (kv: q8_0)7B0/39 (0%)1.6/10 (30 pieces)76 tok/s output
762 tok/s prefill
32K9.6 GB

Earlier round — 2026-08-20

M4 Max · 64 GB · llama-cpp + mlx · 3 trials per task · 2026-08-20 · gezel 3bccbb5d · catalog 0.1.36+local.8c5b93111.qwen3.8-ud — different device, different gezel build, different catalog version

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q3 (kv: q8_0)27B33/33 (100%)8.5/10 (6 pieces)22.9 tok/s output
227 tok/s prefill
256K24.1 GB
qwen3.8-27b-q4 (kv: q8_0)27B33/33 (100%)8.2/10 (9 pieces)22.9 tok/s output
227 tok/s prefill
256K26.9 GB
gemma4-31b-q4 (kv: f16)30.7B32/33 (97%)6.2/10 (9 pieces)22.1 tok/s output
196 tok/s prefill
256K45.8 GB
qwen3.8-27b-q6 (kv: q8_0)27B32/33 (97%)8/10 (7 pieces)18.8 tok/s output
227 tok/s prefill
256K33.2 GB
qwen3.8-27b-q8 (kv: q8_0)27B32/33 (97%)8.3/10 (6 pieces)13.4 tok/s output
219 tok/s prefill
256K38.5 GB
qwen3.6-35b-a3b-q4 (kv: q8_0)35B31/33 (94%)8/10 (9 pieces)75.6 tok/s output
1,156 tok/s prefill
256K25.4 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q3 (kv: q8_0)27B35/39 (90%)7.5/10 (24 pieces)22.9 tok/s output
227 tok/s prefill
256K22.9 GB
qwen3.8-27b-q4 (kv: q8_0)27B35/39 (90%)6.9/10 (25 pieces)22.9 tok/s output
227 tok/s prefill
256K27.1 GB
qwen3.8-27b-q6 (kv: q8_0)27B35/39 (90%)7.1/10 (25 pieces)18.8 tok/s output
227 tok/s prefill
256K32.6 GB
qwen3.6-35b-a3b-q4 (kv: q8_0)35B33/39 (85%)6.9/10 (24 pieces)75.6 tok/s output
1,156 tok/s prefill
256K25.1 GB
qwen3.8-27b-q8 (kv: q8_0)27B33/39 (85%)7.3/10 (24 pieces)13.4 tok/s output
219 tok/s prefill
256K37.6 GB
gemma4-31b-q4 (kv: f16)30.7B28/39 (72%)6.4/10 (19 pieces)22.1 tok/s output
196 tok/s prefill
256K45.7 GB

Earlier round — 2026-08-20

DGX Spark Class · 122 GB · ds4 + llama-cpp · 3 trials per task · 2026-08-20 · gezel 2745e97 · catalog 0.1.36 — different device, different gezel build, different catalog version

Memory isn't shown for this round. This machine's graphics chip shares the computer's memory in a way our measurement can't see, so the number would read far below what the model really uses.

General capability
ModelSizeTasks passedQualityPerformanceContext
deepseek-v4-flash-284b-q2284B33/33 (100%)7.8/10 (7 pieces)16.9 tok/s output
680 tok/s prefill
—
qwen3.8-27b-q4 (kv: q8_0)27B32/32 (100%)8.2/10 (9 pieces)12.1 tok/s output
706 tok/s prefill
256K
qwen3.6-27b-q8 (kv: q8_0)27B30/31 (97%)8.2/10 (9 pieces)7.8 tok/s output
725 tok/s prefill
256K
btl4-35b-q4 (kv: q8_0)35B27/31 (87%)7.8/10 (8 pieces)73.2 tok/s output
1,608 tok/s prefill
64K
gemma4-e4b-q4 (kv: f16)8B25/33 (76%)6.2/10 (3 pieces)65.4 tok/s output
4,698 tok/s prefill
128K
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContext
qwen3.8-27b-q4 (kv: q8_0)27B35/39 (90%)7.2/10 (27 pieces)12.1 tok/s output
706 tok/s prefill
256K
deepseek-v4-flash-284b-q2284B34/39 (87%)6.5/10 (26 pieces)16.9 tok/s output
680 tok/s prefill
—
qwen3.6-27b-q8 (kv: q8_0)27B32/39 (82%)6.8/10 (27 pieces)7.8 tok/s output
725 tok/s prefill
256K
btl4-35b-q4 (kv: q8_0)35B31/39 (79%)6.7/10 (27 pieces)73.2 tok/s output
1,608 tok/s prefill
64K
gemma4-e4b-q4 (kv: f16)8B31/39 (79%)5.2/10 (26 pieces)65.4 tok/s output
4,698 tok/s prefill
128K

Earlier round — 2026-08-14

M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-14 · gezel e367f442 · catalog 0.1.29 — different device, different gezel build, different catalog version

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q427B33/33 (100%)8.2/10 (9 pieces)22.4 tok/s output
233 tok/s prefill
256K27.6 GB
qwen3.6-27b-q427B32/33 (97%)7.9/10 (9 pieces)22.9 tok/s output
230 tok/s prefill
256K29.0 GB
qwen3.6-35b-a3b-q435B31/33 (94%)8/10 (9 pieces)78.5 tok/s output
1,182 tok/s prefill
256K26.0 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q427B35/39 (90%)7.1/10 (26 pieces)22.4 tok/s output
233 tok/s prefill
256K27.3 GB
qwen3.6-35b-a3b-q435B34/39 (87%)7/10 (27 pieces)78.5 tok/s output
1,182 tok/s prefill
256K25.5 GB

Not published — some tasks could not be measured on this round: qwen3.6-27b-q4 (1 task(s) unmeasured)

Earlier round — 2026-08-13

M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-13 · gezel c5904085 · catalog 0.1.23 — different device, different gezel build, different catalog version

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
gemma4-26b-q425.2B33/33 (100%)5.7/10 (9 pieces)107.5 tok/s output
1,264 tok/s prefill
118K41.7 GB
muse-glimmer-30b-q430B30/33 (91%)6.5/10 (9 pieces)26 tok/s output
226 tok/s prefill
128K19.2 GB
qwen3.6-35b-a3b-q435B30/33 (91%)8.1/10 (9 pieces)71.3 tok/s output
1,104 tok/s prefill
256K25.9 GB
nemotron3.5-lightning-30b-q430B26/33 (79%)5.6/10 (9 pieces)93.2 tok/s output
1,030 tok/s prefill
1024K29.4 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.6-35b-a3b-q435B33/39 (85%)6.6/10 (27 pieces)71.3 tok/s output
1,104 tok/s prefill
256K25.1 GB
gemma4-26b-q425.2B28/39 (72%)6/10 (19 pieces)107.5 tok/s output
1,264 tok/s prefill
118K40.8 GB
muse-glimmer-30b-q430B25/39 (64%)5.5/10 (25 pieces)26 tok/s output
228 tok/s prefill
128K19.2 GB
nemotron3.5-lightning-30b-q430B23/39 (59%)5.6/10 (27 pieces)93.2 tok/s output
1,030 tok/s prefill
1024K28.8 GB

Earlier round — 2026-08-11

M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-11 · gezel b2ea2819 · catalog 0.1.20 — different device, different gezel build, different catalog version

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
gemma4-26b-q425.2B33/33 (100%)6.4/10 (9 pieces)89.6 tok/s output
1,093 tok/s prefill
118K40.9 GB
qwen3.6-27b-q427B33/33 (100%)8/10 (9 pieces)18.5 tok/s output
198 tok/s prefill
256K26.4 GB
qwen3.6-35b-a3b-q435B32/33 (97%)8.2/10 (9 pieces)76.4 tok/s output
1,125 tok/s prefill
256K26.5 GB
gemma4-31b-q430.7B31/33 (94%)6.7/10 (9 pieces)23.3 tok/s output
202 tok/s prefill
256K43.1 GB
gemma4-e4b-q48B31/33 (94%)4.7/10 (9 pieces)94.7 tok/s output
1,398 tok/s prefill
128K13.5 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.6-35b-a3b-q435B33/39 (85%)6.6/10 (26 pieces)76.4 tok/s output
1,125 tok/s prefill
256K25.6 GB
gemma4-26b-q425.2B30/39 (77%)6.1/10 (20 pieces)89.6 tok/s output
1,093 tok/s prefill
118K40.2 GB
gemma4-e4b-q48B29/39 (74%)5.4/10 (25 pieces)94.7 tok/s output
1,398 tok/s prefill
128K12.4 GB
qwen3.6-27b-q427B29/39 (74%)5.6/10 (24 pieces)18.5 tok/s output
198 tok/s prefill
256K27.0 GB
gemma4-31b-q430.7B28/39 (72%)6.1/10 (20 pieces)23.3 tok/s output
202 tok/s prefill
256K45.1 GB

Earlier round — 2026-08-09

M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-09 · gezel e2859602 · catalog 0.1.17 — different device, different gezel build, different catalog version

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
gemma4-31b-q430.7B33/33 (100%)6.8/10 (9 pieces)22.3 tok/s output
195 tok/s prefill
64K31.0 GB
gemma4-12b-q412B28/33 (85%)6.9/10 (9 pieces)50.9 tok/s output
543 tok/s prefill
64K30.6 GB
gemma4-e4b-q48B24/33 (73%)4.6/10 (9 pieces)97.8 tok/s output
1,411 tok/s prefill
64K10.4 GB
ornith-9b-q49B24/33 (73%)6.7/10 (9 pieces)64.9 tok/s output
657 tok/s prefill
64K9.8 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
ornith-9b-q49B34/39 (87%)6.5/10 (26 pieces)64.9 tok/s output
657 tok/s prefill
64K8.9 GB
gemma4-12b-q412B29/39 (74%)4.8/10 (26 pieces)50.9 tok/s output
543 tok/s prefill
64K29.7 GB
gemma4-e4b-q48B28/39 (72%)5.2/10 (22 pieces)97.8 tok/s output
1,411 tok/s prefill
64K8.9 GB
gemma4-31b-q430.7B25/39 (64%)6.5/10 (18 pieces)22.3 tok/s output
195 tok/s prefill
64K30.1 GB

Reading the table

Tasks passed counts every attempt across every job in the set. A model with 24/33 (73%) finished 24 of 33 attempts correctly. If any job in the set ran fewer than three times you'll see a raw count instead of a percentage — too small a sample to quote as a rate. A model whose results were incomplete is left out of the table entirely rather than shown with a gap.

Quality is an AI reviewer's opinion of the finished work, and the count beside it is how many pieces that opinion covers. It only grades work that was actually produced, so a model that fails often is judged on its successes alone — 6.5/10 (18 pieces) is a weaker claim than 6.5/10 (26 pieces). Treat it as colour next to the pass rate, never as a substitute for it.

Performance shows two measured speeds on the machine named above: output speed is how fast the model writes its answer, and prefill speed is how fast it takes in your prompt and documents. Both matter for how a gezel feels — prefill governs the pause before it starts, while output governs how fast text appears. A dash means that round did not record a throughput probe.

Context is the working memory the model was given for these runs — how much it can hold at once. Memory used is the most memory the model and its engine held on the test machine. What that number covers depends on the machine, so each round says which one it measured:

  • On a Mac the graphics chip shares the computer's memory and our measurement sees all of it, so Memory used is the model's whole footprint.

  • On a PC with a separate graphics card the column is called System memory. It counts the computer's own memory only, not what the model keeps on the graphics card.

  • On machines like the DGX Spark, where the graphics chip shares memory in a way our measurement can't see, the column is left out. The number would read far below what the model really uses.

Treat these as a rough guide from one machine, not a promise about yours. The Models catalogue in the app is the better check for whether a model fits your computer. When the engine reported its KV-cache precision, it appears beside the model name — for example, (kv: q8_0).

Each test round keeps all of its tables together under one provenance stamp. Earlier rounds stay separate from the latest because a change to gezel or to the task set can move a score without any model changing.

Size is the model's parameter count where we know it. Bigger is often but not always better: on office work in particular, some smaller models beat larger ones, and a model family's habits matter more than its size.

Choosing from this

A high score on the office set is the better guide for everyday document, planning, and analysis work. A high score on the general set matters more if you want a gezel writing or fixing code, and the engineering set, where a round has it, is the stricter test of that.

If a model you're considering isn't listed, it hasn't been measured here yet — which is not a verdict on it either way. The Models catalogue in the app will still tell you whether it fits this machine.

Watch this article as a slideshow