gezel Gezel Handboek

Model scorecard

These are measured results, not estimates. Each number is the count of jobs a model actually finished correctly, checked by a program on a real machine.

How we test models explains what was counted and what these numbers do not tell you. The short version: each job is run three times, a job passes only when every requirement is met, and runs lost to machine trouble are set aside rather than blamed on the model.

Results by test round

Each round has two tables measured on the same machine and build. General capability covers writing a small working program, fixing a bug from its symptoms, following a procedure and stopping at a problem, and turning several documents into one reconciled summary. Office and knowledge work covers a customer notice under a hard word limit, a meeting turned into an action register, a cited research brief, an experiment read-out, a spreadsheet model, a slide deck, and a Word document.

Latest round — 2026-08-30

DGX Spark Class · NVIDIA GB10 · 122 GB · llama-cpp · 3 trials per task · 2026-08-30 · gezel 8d7bb7e5 · catalog 0.1.49

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q4 (kv: f16)27B32/33 (97%)8.2/10 (9 pieces)32.4 tok/s output
777 tok/s prefill
256K6.1 GB
qwen3.6-35b-a3b-q4 (kv: f16)35B31/33 (94%)8/10 (9 pieces)82 tok/s output
1,690 tok/s prefill
64K4.2 GB
ornith1.5-9b-q8 (kv: f16)9B28/32 (88%)7.1/10 (9 pieces)25.6 tok/s output
2,117 tok/s prefill
256K5.2 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q4 (kv: f16)27B35/39 (90%)7.2/10 (27 pieces)32.4 tok/s output
777 tok/s prefill
256K4.9 GB
qwen3.6-35b-a3b-q4 (kv: f16)35B34/39 (87%)6.9/10 (26 pieces)82 tok/s output
1,690 tok/s prefill
64K4.1 GB
ornith1.5-9b-q8 (kv: f16)9B32/39 (82%)6.4/10 (27 pieces)25.6 tok/s output
2,117 tok/s prefill
256K4.5 GB

Earlier round — 2026-08-27

AMD Ryzen 9 7950X3D 16-Core Processor · AMD Radeon AI PRO R9700 · 64 GB · llama-cpp · 3 trials per task · 2026-08-27 · gezel 3cf0b7a2 · catalog 0.1.44 — different device, different gezel build, different catalog version

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q4 (kv: q8_0)27B32/33 (97%)8.1/10 (18 pieces)31.8 tok/s output
681 tok/s prefill
256K15.8 GB
qwen3.6-35b-a3b-q4 (kv: q8_0)35B30/33 (91%)7.8/10 (9 pieces)131.1 tok/s output
2,102 tok/s prefill
64K16.1 GB
gemma4-31b-q4 (kv: f16)30.7B29/33 (88%)7.1/10 (7 pieces)28.7 tok/s output
407 tok/s prefill
99K16.8 GB
gemma4-26b-q4 (kv: f16)25.2B28/33 (85%)6.9/10 (6 pieces)131.2 tok/s output
1,820 tok/s prefill
64K14.9 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q4 (kv: q8_0)27B36/39 (92%)7.5/10 (26 pieces)31.8 tok/s output
681 tok/s prefill
256K14.9 GB
qwen3.6-35b-a3b-q4 (kv: q8_0)35B35/39 (90%)6.8/10 (26 pieces)131.1 tok/s output
2,102 tok/s prefill
64K15.6 GB
gemma4-31b-q4 (kv: f16)30.7B30/39 (77%)6.5/10 (20 pieces)28.7 tok/s output
407 tok/s prefill
99K16.9 GB
gemma4-26b-q4 (kv: f16)25.2B26/39 (67%)5.7/10 (22 pieces)131.2 tok/s output
1,820 tok/s prefill
64K14.7 GB

Earlier round — 2026-08-26

M4 Max · 64 GB · llama-cpp + mlx · 3 trials per task · 2026-08-26 · gezel d1d5e5c0 · catalog 0.1.42 — different device, different gezel build, different catalog version

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q2 (kv: q8_0)27B32/33 (97%)7.9/10 (9 pieces)24.1 tok/s output
241 tok/s prefill
256K21.1 GB
qwen3.8-27b-q3 (kv: q8_0)27B32/33 (97%)8/10 (9 pieces)22.7 tok/s output
239 tok/s prefill
256K24.5 GB
qwen3.8-27b-q4 (kv: q8_0)27B31/33 (94%)7.8/10 (17 pieces)21.9 tok/s output
232 tok/s prefill
256K28.4 GB
qwen3.8-27b-q427B31/33 (94%)7.8/10 (17 pieces)68.4 tok/s output
232 tok/s prefill
qwen3.8-27b-iq1-s (kv: q8_0)27B14/33 (42%)11.1 tok/s output
190 tok/s prefill
256K21.1 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q4 (kv: q8_0)27B36/39 (92%)7.3/10 (53 pieces)21.9 tok/s output
232 tok/s prefill
256K26.2 GB
qwen3.8-27b-q427B36/39 (92%)7.3/10 (53 pieces)68.4 tok/s output
232 tok/s prefill
qwen3.8-27b-q2 (kv: q8_0)27B35/39 (90%)7.3/10 (25 pieces)24.1 tok/s output
241 tok/s prefill
256K19.8 GB
qwen3.8-27b-q3 (kv: q8_0)27B34/39 (87%)7.1/10 (27 pieces)22.7 tok/s output
239 tok/s prefill
256K24.4 GB
qwen3.8-27b-iq1-s (kv: q8_0)27B14/39 (36%)3.9/10 (25 pieces)11.1 tok/s output
190 tok/s prefill
256K20.2 GB

Earlier round — 2026-08-22

M4 Max · 64 GB · llama-cpp + mlx · 3 trials per task · 2026-08-22 · gezel b3be0a7f · catalog 0.1.39 — different device, different gezel build, different catalog version

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
gemma4-12b-q4 (kv: f16)12B31/33 (94%)6.9/10 (6 pieces)48.2 tok/s output
508 tok/s prefill
99K42.2 GB
gemma4-12b-q8 (kv: f16)12B30/33 (91%)8.1/10 (3 pieces)30.3 tok/s output
506 tok/s prefill
83K42.0 GB
btl4-compact-iq2 (kv: q8_0)35B29/33 (88%)7.1/10 (9 pieces)80.4 tok/s output
1,120 tok/s prefill
256K14.6 GB
ornith1.5-9b-q8 (kv: q8_0)9B28/33 (85%)4.4/10 (5 pieces)47 tok/s output
823 tok/s prefill
256K15.9 GB
qwen3.5-9b-q4 (kv: q8_0)9B28/33 (85%)5.3/10 (9 pieces)62.9 tok/s output
787 tok/s prefill
256K13.1 GB
gemma4-e4b-q4 (kv: f16)8B26/33 (79%)5.4/10 (9 pieces)96.5 tok/s output
1,442 tok/s prefill
128K13.7 GB
qwen3.5-4b-q4 (kv: q8_0)4B22/33 (67%)6.6/10 (9 pieces)89.3 tok/s output
1,336 tok/s prefill
256K11.5 GB
gemma4-e2b-q4 (kv: f16)2.3B17/33 (52%)4.3/10 (9 pieces)144.1 tok/s output
2,574 tok/s prefill
128K6.9 GB
lfm2.5-2.6b-q4 (kv: q8_0)2.6B13/33 (39%)3.6/10 (8 pieces)168.4 tok/s output
2,125 tok/s prefill
125K4.9 GB
qwen3.5-2b-q4 (kv: q8_0)2B9/33 (27%)3.2/10 (9 pieces)155.3 tok/s output
3,084 tok/s prefill
256K5.2 GB
mistral-7b-q4 (kv: q8_0)7B4/33 (12%)76 tok/s output
762 tok/s prefill
32K9.6 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
gemma4-12b-q8 (kv: f16)12B35/39 (90%)5.8/10 (24 pieces)30.3 tok/s output
506 tok/s prefill
83K41.8 GB
gemma4-12b-q4 (kv: f16)12B34/39 (87%)5.8/10 (24 pieces)48.2 tok/s output
508 tok/s prefill
99K41.5 GB
ornith1.5-9b-q8 (kv: q8_0)9B34/39 (87%)6.5/10 (27 pieces)47 tok/s output
823 tok/s prefill
256K14.9 GB
gemma4-e4b-q4 (kv: f16)8B31/39 (79%)5.5/10 (22 pieces)96.5 tok/s output
1,442 tok/s prefill
128K12.5 GB
qwen3.5-9b-q4 (kv: q8_0)9B31/39 (79%)6.2/10 (25 pieces)62.9 tok/s output
787 tok/s prefill
256K11.3 GB
qwen3.5-4b-q4 (kv: q8_0)4B30/39 (77%)6.2/10 (24 pieces)89.3 tok/s output
1,336 tok/s prefill
256K9.9 GB
btl4-compact-iq2 (kv: q8_0)35B27/39 (69%)6.2/10 (25 pieces)80.4 tok/s output
1,120 tok/s prefill
256K14.1 GB
gemma4-e2b-q4 (kv: f16)2.3B16/39 (41%)4.5/10 (20 pieces)144.1 tok/s output
2,574 tok/s prefill
128K6.2 GB
lfm2.5-2.6b-q4 (kv: q8_0)2.6B13/39 (33%)4.4/10 (27 pieces)168.4 tok/s output
2,125 tok/s prefill
125K4.8 GB
qwen3.5-2b-q4 (kv: q8_0)2B10/39 (26%)4.1/10 (26 pieces)155.3 tok/s output
3,084 tok/s prefill
256K4.6 GB
mistral-7b-q4 (kv: q8_0)7B0/39 (0%)1.6/10 (30 pieces)76 tok/s output
762 tok/s prefill
32K9.6 GB

Earlier round — 2026-08-20

M4 Max · 64 GB · llama-cpp + mlx · 3 trials per task · 2026-08-20 · gezel 3bccbb5d · catalog 0.1.36+local.8c5b93111.qwen3.8-ud — different device, different gezel build, different catalog version

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q3 (kv: q8_0)27B33/33 (100%)8.5/10 (6 pieces)22.9 tok/s output
227 tok/s prefill
256K24.1 GB
qwen3.8-27b-q4 (kv: q8_0)27B33/33 (100%)8.2/10 (9 pieces)22.9 tok/s output
227 tok/s prefill
256K26.9 GB
gemma4-31b-q4 (kv: f16)30.7B32/33 (97%)6.2/10 (9 pieces)22.1 tok/s output
196 tok/s prefill
256K45.8 GB
qwen3.8-27b-q6 (kv: q8_0)27B32/33 (97%)8/10 (7 pieces)18.8 tok/s output
227 tok/s prefill
256K33.2 GB
qwen3.8-27b-q8 (kv: q8_0)27B32/33 (97%)8.3/10 (6 pieces)13.4 tok/s output
219 tok/s prefill
256K38.5 GB
qwen3.6-35b-a3b-q4 (kv: q8_0)35B31/33 (94%)8/10 (9 pieces)75.6 tok/s output
1,156 tok/s prefill
256K25.4 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q3 (kv: q8_0)27B35/39 (90%)7.5/10 (24 pieces)22.9 tok/s output
227 tok/s prefill
256K22.9 GB
qwen3.8-27b-q4 (kv: q8_0)27B35/39 (90%)6.9/10 (25 pieces)22.9 tok/s output
227 tok/s prefill
256K27.1 GB
qwen3.8-27b-q6 (kv: q8_0)27B35/39 (90%)7.1/10 (25 pieces)18.8 tok/s output
227 tok/s prefill
256K32.6 GB
qwen3.6-35b-a3b-q4 (kv: q8_0)35B33/39 (85%)6.9/10 (24 pieces)75.6 tok/s output
1,156 tok/s prefill
256K25.1 GB
qwen3.8-27b-q8 (kv: q8_0)27B33/39 (85%)7.3/10 (24 pieces)13.4 tok/s output
219 tok/s prefill
256K37.6 GB
gemma4-31b-q4 (kv: f16)30.7B28/39 (72%)6.4/10 (19 pieces)22.1 tok/s output
196 tok/s prefill
256K45.7 GB

Earlier round — 2026-08-20

DGX Spark Class · 122 GB · ds4 + llama-cpp · 3 trials per task · 2026-08-20 · gezel 2745e97 · catalog 0.1.36 — different device, different gezel build, different catalog version

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
deepseek-v4-flash-284b-q2284B33/33 (100%)7.8/10 (7 pieces)16.9 tok/s output
680 tok/s prefill
qwen3.8-27b-q4 (kv: q8_0)27B32/32 (100%)8.2/10 (9 pieces)12.1 tok/s output
706 tok/s prefill
256K18.0 GB
qwen3.6-27b-q8 (kv: q8_0)27B30/31 (97%)8.2/10 (9 pieces)7.8 tok/s output
725 tok/s prefill
256K29.3 GB
btl4-35b-q4 (kv: q8_0)35B27/31 (87%)7.8/10 (8 pieces)73.2 tok/s output
1,608 tok/s prefill
64K22.2 GB
gemma4-e4b-q4 (kv: f16)8B25/33 (76%)6.2/10 (3 pieces)65.4 tok/s output
4,698 tok/s prefill
128K5.7 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q4 (kv: q8_0)27B35/39 (90%)7.2/10 (27 pieces)12.1 tok/s output
706 tok/s prefill
256K5.4 GB
deepseek-v4-flash-284b-q2284B34/39 (87%)6.5/10 (26 pieces)16.9 tok/s output
680 tok/s prefill
qwen3.6-27b-q8 (kv: q8_0)27B32/39 (82%)6.8/10 (27 pieces)7.8 tok/s output
725 tok/s prefill
256K29.2 GB
btl4-35b-q4 (kv: q8_0)35B31/39 (79%)6.7/10 (27 pieces)73.2 tok/s output
1,608 tok/s prefill
64K4.8 GB
gemma4-e4b-q4 (kv: f16)8B31/39 (79%)5.2/10 (26 pieces)65.4 tok/s output
4,698 tok/s prefill
128K5.1 GB

Earlier round — 2026-08-14

M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-14 · gezel e367f442 · catalog 0.1.29 — different device, different gezel build, different catalog version

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q427B33/33 (100%)8.2/10 (9 pieces)22.4 tok/s output
233 tok/s prefill
256K27.6 GB
qwen3.6-27b-q427B32/33 (97%)7.9/10 (9 pieces)22.9 tok/s output
230 tok/s prefill
256K29.0 GB
qwen3.6-35b-a3b-q435B31/33 (94%)8/10 (9 pieces)78.5 tok/s output
1,182 tok/s prefill
256K26.0 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.8-27b-q427B35/39 (90%)7.1/10 (26 pieces)22.4 tok/s output
233 tok/s prefill
256K27.3 GB
qwen3.6-35b-a3b-q435B34/39 (87%)7/10 (27 pieces)78.5 tok/s output
1,182 tok/s prefill
256K25.5 GB

Not published — some tasks could not be measured on this round: qwen3.6-27b-q4 (1 task(s) unmeasured)

Earlier round — 2026-08-13

M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-13 · gezel c5904085 · catalog 0.1.23 — different device, different gezel build, different catalog version

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
gemma4-26b-q425.2B33/33 (100%)5.7/10 (9 pieces)107.5 tok/s output
1,264 tok/s prefill
118K41.7 GB
muse-glimmer-30b-q430B30/33 (91%)6.5/10 (9 pieces)26 tok/s output
226 tok/s prefill
128K19.2 GB
qwen3.6-35b-a3b-q435B30/33 (91%)8.1/10 (9 pieces)71.3 tok/s output
1,104 tok/s prefill
256K25.9 GB
nemotron3.5-lightning-30b-q430B26/33 (79%)5.6/10 (9 pieces)93.2 tok/s output
1,030 tok/s prefill
1024K29.4 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.6-35b-a3b-q435B33/39 (85%)6.6/10 (27 pieces)71.3 tok/s output
1,104 tok/s prefill
256K25.1 GB
gemma4-26b-q425.2B28/39 (72%)6/10 (19 pieces)107.5 tok/s output
1,264 tok/s prefill
118K40.8 GB
muse-glimmer-30b-q430B25/39 (64%)5.5/10 (25 pieces)26 tok/s output
228 tok/s prefill
128K19.2 GB
nemotron3.5-lightning-30b-q430B23/39 (59%)5.6/10 (27 pieces)93.2 tok/s output
1,030 tok/s prefill
1024K28.8 GB

Earlier round — 2026-08-11

M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-11 · gezel b2ea2819 · catalog 0.1.20 — different device, different gezel build, different catalog version

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
gemma4-26b-q425.2B33/33 (100%)6.4/10 (9 pieces)89.6 tok/s output
1,093 tok/s prefill
118K40.9 GB
qwen3.6-27b-q427B33/33 (100%)8/10 (9 pieces)18.5 tok/s output
198 tok/s prefill
256K26.4 GB
qwen3.6-35b-a3b-q435B32/33 (97%)8.2/10 (9 pieces)76.4 tok/s output
1,125 tok/s prefill
256K26.5 GB
gemma4-31b-q430.7B31/33 (94%)6.7/10 (9 pieces)23.3 tok/s output
202 tok/s prefill
256K43.1 GB
gemma4-e4b-q48B31/33 (94%)4.7/10 (9 pieces)94.7 tok/s output
1,398 tok/s prefill
128K13.5 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
qwen3.6-35b-a3b-q435B33/39 (85%)6.6/10 (26 pieces)76.4 tok/s output
1,125 tok/s prefill
256K25.6 GB
gemma4-26b-q425.2B30/39 (77%)6.1/10 (20 pieces)89.6 tok/s output
1,093 tok/s prefill
118K40.2 GB
gemma4-e4b-q48B29/39 (74%)5.4/10 (25 pieces)94.7 tok/s output
1,398 tok/s prefill
128K12.4 GB
qwen3.6-27b-q427B29/39 (74%)5.6/10 (24 pieces)18.5 tok/s output
198 tok/s prefill
256K27.0 GB
gemma4-31b-q430.7B28/39 (72%)6.1/10 (20 pieces)23.3 tok/s output
202 tok/s prefill
256K45.1 GB

Earlier round — 2026-08-09

M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-09 · gezel e2859602 · catalog 0.1.17 — different device, different gezel build, different catalog version

General capability
ModelSizeTasks passedQualityPerformanceContextMemory used
gemma4-31b-q430.7B33/33 (100%)6.8/10 (9 pieces)22.3 tok/s output
195 tok/s prefill
64K31.0 GB
gemma4-12b-q412B28/33 (85%)6.9/10 (9 pieces)50.9 tok/s output
543 tok/s prefill
64K30.6 GB
gemma4-e4b-q48B24/33 (73%)4.6/10 (9 pieces)97.8 tok/s output
1,411 tok/s prefill
64K10.4 GB
ornith-9b-q49B24/33 (73%)6.7/10 (9 pieces)64.9 tok/s output
657 tok/s prefill
64K9.8 GB
Office and knowledge work
ModelSizeTasks passedQualityPerformanceContextMemory used
ornith-9b-q49B34/39 (87%)6.5/10 (26 pieces)64.9 tok/s output
657 tok/s prefill
64K8.9 GB
gemma4-12b-q412B29/39 (74%)4.8/10 (26 pieces)50.9 tok/s output
543 tok/s prefill
64K29.7 GB
gemma4-e4b-q48B28/39 (72%)5.2/10 (22 pieces)97.8 tok/s output
1,411 tok/s prefill
64K8.9 GB
gemma4-31b-q430.7B25/39 (64%)6.5/10 (18 pieces)22.3 tok/s output
195 tok/s prefill
64K30.1 GB

Reading the table

Tasks passed counts every attempt across every job in the set. A model with 24/33 (73%) finished 24 of 33 attempts correctly. If any job in the set ran fewer than three times you'll see a raw count instead of a percentage — too small a sample to quote as a rate. A model whose results were incomplete is left out of the table entirely rather than shown with a gap.

Quality is an AI reviewer's opinion of the finished work, and the count beside it is how many pieces that opinion covers. It only grades work that was actually produced, so a model that fails often is judged on its successes alone — 6.5/10 (18 pieces) is a weaker claim than 6.5/10 (26 pieces). Treat it as colour next to the pass rate, never as a substitute for it.

Performance shows two measured speeds on the machine named above: output speed is how fast the model writes its answer, and prefill speed is how fast it takes in your prompt and documents. Both matter for how a gezel feels — prefill governs the pause before it starts, while output governs how fast text appears. A dash means that round did not record a throughput probe.

Context is the working memory the model was given for these runs — how much it can hold at once. Memory used is the peak RAM the model and its engine actually occupied, which is the number to check against your own machine. When the engine reported its KV-cache precision, it appears beside the model name — for example, (kv: q8_0).

Each test round keeps its General capability and Office and knowledge work tables together under one provenance stamp. Earlier rounds stay separate from the latest because a change to gezel or to the task set can move a score without any model changing.

Size is the model's parameter count where we know it. Bigger is often but not always better: on office work in particular, some smaller models beat larger ones, and a model family's habits matter more than its size.

Choosing from this

A high score on the office set is the better guide for everyday document, planning, and analysis work. A high score on the general set matters more if you want a gezel writing or fixing code.

If a model you're considering isn't listed, it hasn't been measured here yet — which is not a verdict on it either way. The Models catalogue in the app will still tell you whether it fits this machine.

Watch this article as a slideshow