gezel Gezel Handboek

Model scorecard

These are measured results, not estimates. Each number is the count of jobs a model actually finished correctly, checked by a program on a real machine.

How we test models explains what was counted and what these numbers do not tell you. The short version: each job is run three times, a job passes only when every requirement is met, and runs lost to machine trouble are set aside rather than blamed on the model.

General capability

Writing a small working program, fixing a bug from its symptoms, following a procedure and stopping at a problem, turning several documents into one reconciled summary.

M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-14 · gezel e367f442 · catalog 0.1.29

ModelSizeTasks passedQualityReads atWrites atContextMemory used
qwen3.8-
27b-q4
27B33/33 (100%)8.2/10 (9 pieces)233 tok/s22.4 tok/s256K27.6 GB
qwen3.6-
27b-q4
27B32/33 (97%)7.9/10 (9 pieces)230 tok/s22.9 tok/s256K29.0 GB
qwen3.6-
35b-a3b-q4
35B31/33 (94%)8/10 (9 pieces)1,182 tok/s78.5 tok/s256K26.0 GB

Earlier round — 2026-08-13

M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-13 · gezel c5904085 · catalog 0.1.23 — different gezel build, different catalog version

ModelSizeTasks passedQualityReads atWrites atContextMemory used
gemma4-
26b-q4
25.2B33/33 (100%)5.7/10 (9 pieces)1,264 tok/s107.5 tok/s118K41.7 GB
muse-glimmer-
30b-q4
30B30/33 (91%)6.5/10 (9 pieces)226 tok/s26 tok/s128K19.2 GB
qwen3.6-
35b-a3b-q4
35B30/33 (91%)8.1/10 (9 pieces)1,104 tok/s71.3 tok/s256K25.9 GB
nemotron3.5-lightning-
30b-q4
30B26/33 (79%)5.6/10 (9 pieces)1,030 tok/s93.2 tok/s1024K29.4 GB

Earlier round — 2026-08-11

M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-11 · gezel b2ea2819 · catalog 0.1.20 — different gezel build, different catalog version

ModelSizeTasks passedQualityReads atWrites atContextMemory used
gemma4-
26b-q4
25.2B33/33 (100%)6.4/10 (9 pieces)1,093 tok/s89.6 tok/s118K40.9 GB
qwen3.6-
27b-q4
27B33/33 (100%)8/10 (9 pieces)198 tok/s18.5 tok/s256K26.4 GB
qwen3.6-
35b-a3b-q4
35B32/33 (97%)8.2/10 (9 pieces)1,125 tok/s76.4 tok/s256K26.5 GB
gemma4-
31b-q4
30.7B31/33 (94%)6.7/10 (9 pieces)202 tok/s23.3 tok/s256K43.1 GB
gemma4-e4b-q48B31/33 (94%)4.7/10 (9 pieces)1,398 tok/s94.7 tok/s128K13.5 GB

Office and knowledge work

A customer notice under a hard word limit, a meeting turned into an action register, a cited research brief, an experiment read-out, a spreadsheet model, a slide deck, a Word document.

M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-14 · gezel e367f442 · catalog 0.1.29

ModelSizeTasks passedQualityReads atWrites atContextMemory used
qwen3.8-
27b-q4
27B35/39 (90%)7.1/10 (26 pieces)233 tok/s22.4 tok/s256K27.3 GB
qwen3.6-
35b-a3b-q4
35B34/39 (87%)7/10 (27 pieces)1,182 tok/s78.5 tok/s256K25.5 GB

Not published — some tasks could not be measured on this round:

  • qwen3.6-27b-q4: 1 task(s) unmeasured

Earlier round — 2026-08-13

M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-13 · gezel c5904085 · catalog 0.1.23 — different gezel build, different catalog version

ModelSizeTasks passedQualityReads atWrites atContextMemory used
qwen3.6-
35b-a3b-q4
35B33/39 (85%)6.6/10 (27 pieces)1,104 tok/s71.3 tok/s256K25.1 GB
gemma4-
26b-q4
25.2B28/39 (72%)6/10 (19 pieces)1,264 tok/s107.5 tok/s118K40.8 GB
muse-glimmer-
30b-q4
30B25/39 (64%)5.5/10 (25 pieces)228 tok/s26 tok/s128K19.2 GB
nemotron3.5-lightning-
30b-q4
30B23/39 (59%)5.6/10 (27 pieces)1,030 tok/s93.2 tok/s1024K28.8 GB

Earlier round — 2026-08-11

M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-11 · gezel b2ea2819 · catalog 0.1.20 — different gezel build, different catalog version

ModelSizeTasks passedQualityReads atWrites atContextMemory used
qwen3.6-
35b-a3b-q4
35B33/39 (85%)6.6/10 (26 pieces)1,125 tok/s76.4 tok/s256K25.6 GB
gemma4-
26b-q4
25.2B30/39 (77%)6.1/10 (20 pieces)1,093 tok/s89.6 tok/s118K40.2 GB
gemma4-e4b-q48B29/39 (74%)5.4/10 (25 pieces)1,398 tok/s94.7 tok/s128K12.4 GB
qwen3.6-
27b-q4
27B29/39 (74%)5.6/10 (24 pieces)198 tok/s18.5 tok/s256K27.0 GB
gemma4-
31b-q4
30.7B28/39 (72%)6.1/10 (20 pieces)202 tok/s23.3 tok/s256K45.1 GB

Reading the table

Tasks passed counts every attempt across every job in the set. A model with 24/33 (73%) finished 24 of 33 attempts correctly. If any job in the set ran fewer than three times you'll see a raw count instead of a percentage — too small a sample to quote as a rate. A model whose results were incomplete is left out of the table entirely rather than shown with a gap.

Quality is an AI reviewer's opinion of the finished work, and the count beside it is how many pieces that opinion covers. It only grades work that was actually produced, so a model that fails often is judged on its successes alone — 6.5/10 (18 pieces) is a weaker claim than 6.5/10 (26 pieces). Treat it as colour next to the pass rate, never as a substitute for it.

Reads at / Writes at are measured speeds on the machine named above: how fast the model takes in your documents, and how fast it writes its answer. Both matter for how a gezel feels — reading speed governs the pause before it starts, writing speed governs how fast text appears.

Context is the working memory the model was given for these runs — how much it can hold at once. Memory used is the peak RAM the model and its engine actually occupied, which is the number to check against your own machine.

Earlier rounds appear as separate tables below each set, with their own stamps. They are kept apart rather than merged because a change to gezel or to the task set can move a score without any model changing.

Size is the model's parameter count where we know it. Bigger is often but not always better: on office work in particular, some smaller models beat larger ones, and a model family's habits matter more than its size.

Choosing from this

A high score on the office set is the better guide for everyday document, planning, and analysis work. A high score on the general set matters more if you want a gezel writing or fixing code.

If a model you're considering isn't listed, it hasn't been measured here yet — which is not a verdict on it either way. The Models catalogue in the app will still tell you whether it fits this machine.

Watch this article as a slideshow