Model scorecard
These are measured results, not estimates. Each number is the count of jobs a model actually finished correctly, checked by a program on a real machine.
How we test models explains what was counted and what these numbers do not tell you. The short version: each job is run three times, a job passes only when every requirement is met, and runs lost to machine trouble are set aside rather than blamed on the model.
Results by test round
Each round has one table for each task set it ran, all measured on the same machine and build. Most rounds ran two sets. General capability covers writing a small working program, fixing a bug from its symptoms, following a procedure and stopping at a problem, and turning several documents into one reconciled summary. Office and knowledge work covers a customer notice under a hard word limit, a meeting turned into an action register, a cited research brief, an experiment read-out, a spreadsheet model, a slide deck, and a Word document. Some rounds add the two harder sets: Engineering work, which is developer work such as reviewing a change and proving a fix with tests, and Complex workflows, which asks a model to write and repair reusable recipes of its own.
Latest round — 2026-09-18
DGX Spark Class · NVIDIA GB10 · 122 GB · ds4 + llama-cpp · 3 trials per task · 2026-09-18 · gezel bcc9d87c · catalog 0.1.67
Memory isn't shown for this round. This machine's graphics chip shares the computer's memory in a way our measurement can't see, so the number would read far below what the model really uses.
General capability
| Model | Size | Tasks passed | Performance | Context |
|---|---|---|---|---|
| qwen3.8-flash-next-q2 | 180B | 32/33 (97%) | 27.1 tok/s output 433 tok/s prefill | — |
| qwen3.8-flash-next-iq4 (kv: q4_0) | 180B | 31/33 (94%) | 25.8 tok/s output 480 tok/s prefill | 64K |
Not published — some tasks could not be measured on this round: qwen3.8-27b-q4 (1 task(s) unmeasured)
Earlier round — 2026-09-08
DGX Spark Class · NVIDIA GB10 · 122 GB · llama-cpp · 3 trials per task · 2026-09-08 · gezel 35c47b9f · catalog 0.1.56 — different gezel build, different catalog version
Memory isn't shown for this round. This machine's graphics chip shares the computer's memory in a way our measurement can't see, so the number would read far below what the model really uses.
General capability
| Model | Size | Tasks passed | Quality | Performance | Context |
|---|---|---|---|---|---|
| qwen3.8-flash-next-iq4 (kv: q4_0) | 180B | 33/33 (100%) | 8.2/10 (9 pieces) | 24.5 tok/s output 380 tok/s prefill | 64K |
| gemma4-31b-q4 (kv: f16) | 30.7B | 32/33 (97%) | 6.6/10 (9 pieces) | 35.9 tok/s output 484 tok/s prefill | 90K |
| qwen3.8-27b-q4 (kv: f16) | 27B | 32/33 (97%) | 8.1/10 (8 pieces) | 28.6 tok/s output 752 tok/s prefill | 256K |
| qwen3.6-35b-a3b-q4 (kv: f16) | 35B | 31/33 (94%) | 7.8/10 (9 pieces) | 87.4 tok/s output 1,675 tok/s prefill | 64K |
| muse-glimmer-30b-q4 (kv: f16) | 30B | 27/33 (82%) | 6.8/10 (9 pieces) | 12.4 tok/s output 932 tok/s prefill | 128K |
| nemotron3.5-lightning-30b-q4 (kv: f16) | 30B | 22/33 (67%) | 5.1/10 (8 pieces) | 103.2 tok/s output 1,705 tok/s prefill | 64K |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context |
|---|---|---|---|---|---|
| qwen3.6-35b-a3b-q4 (kv: f16) | 35B | 36/39 (92%) | 7.5/10 (26 pieces) | 87.4 tok/s output 1,675 tok/s prefill | 64K |
| qwen3.8-27b-q4 (kv: f16) | 27B | 36/39 (92%) | 7.7/10 (26 pieces) | 28.6 tok/s output 752 tok/s prefill | 256K |
| gemma4-31b-q4 (kv: f16) | 30.7B | 35/39 (90%) | 6.9/10 (25 pieces) | 35.9 tok/s output 484 tok/s prefill | 90K |
| qwen3.8-flash-next-iq4 (kv: q4_0) | 180B | 35/39 (90%) | 8/10 (27 pieces) | 24.5 tok/s output 380 tok/s prefill | 64K |
| muse-glimmer-30b-q4 (kv: f16) | 30B | 28/39 (72%) | 7.6/10 (26 pieces) | 12.4 tok/s output 932 tok/s prefill | 128K |
| nemotron3.5-lightning-30b-q4 (kv: f16) | 30B | 19/39 (49%) | 6.3/10 (26 pieces) | 103.2 tok/s output 1,705 tok/s prefill | 64K |
Engineering work
| Model | Size | Tasks passed | Quality | Performance | Context |
|---|---|---|---|---|---|
| gemma4-31b-q4 (kv: f16) | 30.7B | 20/30 (67%) | 5.1/10 (15 pieces) | 35.9 tok/s output 484 tok/s prefill | 90K |
| qwen3.6-35b-a3b-q4 (kv: f16) | 35B | 17/27 (63%) | 6.7/10 (16 pieces) | 87.4 tok/s output 1,675 tok/s prefill | 64K |
| qwen3.8-flash-next-iq4 (kv: q4_0) | 180B | 15/26 (58%) | 7.7/10 (14 pieces) | 24.3 tok/s output 420 tok/s prefill | 64K |
| qwen3.8-27b-q4 (kv: f16) | 27B | 16/29 (55%) | 7.6/10 (16 pieces) | 28 tok/s output 749 tok/s prefill | 256K |
| muse-glimmer-30b-q4 (kv: f16) | 30B | 13/30 (43%) | 5.7/10 (12 pieces) | 12.3 tok/s output 930 tok/s prefill | 128K |
| nemotron3.5-lightning-30b-q4 (kv: f16) | 30B | 13/30 (43%) | 5/10 (10 pieces) | 103.2 tok/s output 1,705 tok/s prefill | 64K |
Complex workflows
| Model | Size | Tasks passed | Quality | Performance | Context |
|---|---|---|---|---|---|
| gemma4-31b-q4 (kv: f16) | 30.7B | 17/27 (63%) | 2.7/10 (6 pieces) | 35.9 tok/s output 484 tok/s prefill | 90K |
| qwen3.8-flash-next-iq4 (kv: q4_0) | 180B | 16/27 (59%) | 0.5/10 (6 pieces) | 24.3 tok/s output 416 tok/s prefill | 64K |
| qwen3.6-35b-a3b-q4 (kv: f16) | 35B | 15/27 (56%) | 3/10 (6 pieces) | 87.4 tok/s output 1,675 tok/s prefill | 64K |
| qwen3.8-27b-q4 (kv: f16) | 27B | 14/27 (52%) | 2.2/10 (6 pieces) | 28.5 tok/s output 753 tok/s prefill | 256K |
| muse-glimmer-30b-q4 (kv: f16) | 30B | 13/27 (48%) | 0.5/10 (6 pieces) | 12.3 tok/s output 921 tok/s prefill | 128K |
| nemotron3.5-lightning-30b-q4 (kv: f16) | 30B | 7/27 (26%) | 1.9/10 (6 pieces) | 103.2 tok/s output 1,705 tok/s prefill | 64K |
Earlier round — 2026-09-04
AMD Ryzen 9 7950X3D 16-Core Processor · AMD Radeon AI PRO R9700 · 64 GB · llama-cpp · 3 trials per task · 2026-09-04 · gezel f115c698 · catalog 0.1.55 — different device, different gezel build, different catalog version
System memory counts this computer's own memory only. What the model keeps on the separate graphics card isn't included.
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | System memory |
|---|---|---|---|---|---|---|
| qwen3.5-27b-q4 (kv: f16) | 27B | 30/33 (91%) | 6.4/10 (9 pieces) | 55.9 tok/s output 656 tok/s prefill | 159K | 22.2 GB |
| qwen3.5-9b-q4 (kv: f16) | 9B | 26/33 (79%) | 5.3/10 (9 pieces) | 122.9 tok/s output 2,143 tok/s prefill | 256K | 8.9 GB |
| qwen3.5-4b-q4 (kv: f16) | 4B | 16/33 (48%) | 5.3/10 (9 pieces) | 153 tok/s output 3,349 tok/s prefill | 256K | 7.2 GB |
| qwen3.5-2b-q4 (kv: f16) | 2B | 6/33 (18%) | 2.7/10 (9 pieces) | 203.3 tok/s output 7,605 tok/s prefill | 256K | 5.1 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | System memory |
|---|---|---|---|---|---|---|
| qwen3.5-27b-q4 (kv: f16) | 27B | 36/39 (92%) | 6.9/10 (27 pieces) | 55.9 tok/s output 656 tok/s prefill | 159K | 21.1 GB |
| qwen3.5-9b-q4 (kv: f16) | 9B | 32/39 (82%) | 6/10 (27 pieces) | 122.9 tok/s output 2,143 tok/s prefill | 256K | 7.6 GB |
| qwen3.5-4b-q4 (kv: f16) | 4B | 28/39 (72%) | 5.3/10 (26 pieces) | 157 tok/s output 3,429 tok/s prefill | 256K | 6.0 GB |
| qwen3.5-2b-q4 (kv: f16) | 2B | 20/39 (51%) | 4.3/10 (27 pieces) | 203.3 tok/s output 7,605 tok/s prefill | 256K | 3.7 GB |
Earlier round — 2026-09-03
M2 · Apple M2 (integrated GPU) · 16 GB · llama-cpp · 1 trial per task · 2026-09-03 · gezel a06b1070 · catalog 0.1.52 — different device, different gezel build, different catalog version, different trial count
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| gemma4-12b-q4 (kv: f16) | 12B | 10/11 (some tasks run once — count not rate) | 6.6/10 (3 pieces) | 13.7 tok/s output 100 tok/s prefill | 147K | 10.1 GB |
| ornith1.5-9b-q4 (kv: f16) | 9B | 10/11 (some tasks run once — count not rate) | 6.5/10 (1 pieces) | 9.8 tok/s output 154 tok/s prefill | 132K | 10.1 GB |
| qwen3.5-9b-q4 (kv: f16) | 9B | 9/11 (some tasks run once — count not rate) | 3.8/10 (1 pieces) | 9.2 tok/s output 121 tok/s prefill | 126K | 10.8 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| ornith1.5-9b-q4 (kv: f16) | 9B | 12/13 (some tasks run once — count not rate) | 5.8/10 (9 pieces) | 9.8 tok/s output 154 tok/s prefill | 132K | 10.0 GB |
| qwen3.5-9b-q4 (kv: f16) | 9B | 11/13 (some tasks run once — count not rate) | 5.3/10 (9 pieces) | 9.2 tok/s output 121 tok/s prefill | 126K | 10.8 GB |
| gemma4-12b-q4 (kv: f16) | 12B | 10/13 (some tasks run once — count not rate) | 5/10 (9 pieces) | 13.7 tok/s output 100 tok/s prefill | 147K | 10.0 GB |
Earlier round — 2026-09-02
AMD Ryzen AI 9 HX 370 w/ Radeon 890M · NVIDIA GeForce RTX 5070 Ti Laptop GPU · 31 GB · llama-cpp · 1 trial per task · 2026-09-02 · gezel 7694351b · catalog 0.1.52 — different device, different gezel build, different catalog version, different trial count
System memory counts this computer's own memory only. What the model keeps on the separate graphics card isn't included.
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | System memory |
|---|---|---|---|---|---|---|
| gemma4-12b-q4 (kv: f16) | 12B | 10/11 (some tasks run once — count not rate) | 7.4/10 (3 pieces) | 24.6 tok/s output 1,495 tok/s prefill | 154K | 11.2 GB |
| qwen3.6-35b-a3b-q4 (kv: f16) | 35B | 10/11 (some tasks run once — count not rate) | 8.1/10 (3 pieces) | 15.1 tok/s output 80 tok/s prefill | 64K | 21.0 GB |
| ornith1.5-9b-q4 (kv: f16) | 9B | 9/11 (some tasks run once — count not rate) | 7.2/10 (3 pieces) | 57.7 tok/s output 2,230 tok/s prefill | 136K | 7.9 GB |
| qwen3.5-9b-q4 (kv: f16) | 9B | 9/11 (some tasks run once — count not rate) | 5.2/10 (3 pieces) | 29.9 tok/s output 329 tok/s prefill | 130K | 9.8 GB |
| qwen3.8-27b-q2 (kv: f16) | 27B | 8/11 (some tasks run once — count not rate) | 7.9/10 (3 pieces) | 7 tok/s output 55 tok/s prefill | 64K | 14.9 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | System memory |
|---|---|---|---|---|---|---|
| qwen3.5-9b-q4 (kv: f16) | 9B | 12/13 (some tasks run once — count not rate) | 6/10 (9 pieces) | 29.9 tok/s output 329 tok/s prefill | 130K | 8.3 GB |
| qwen3.6-35b-a3b-q4 (kv: f16) | 35B | 12/13 (some tasks run once — count not rate) | 6.9/10 (9 pieces) | 15.1 tok/s output 80 tok/s prefill | 64K | 21.1 GB |
| gemma4-12b-q4 (kv: f16) | 12B | 11/13 (some tasks run once — count not rate) | 5.2/10 (9 pieces) | 24.6 tok/s output 1,495 tok/s prefill | 154K | 10.9 GB |
| ornith1.5-9b-q4 (kv: f16) | 9B | 11/13 (some tasks run once — count not rate) | 5.8/10 (9 pieces) | 57.7 tok/s output 2,230 tok/s prefill | 136K | 7.5 GB |
| qwen3.8-27b-q2 (kv: f16) | 27B | 10/13 (some tasks run once — count not rate) | 6.8/10 (9 pieces) | 7 tok/s output 55 tok/s prefill | 64K | 14.5 GB |
Earlier round — 2026-08-31
DGX Spark Class · NVIDIA GB10 · 122 GB · ds4 + llama-cpp · 3 trials per task · 2026-08-31 · gezel f59825ea · catalog 0.1.51 — different gezel build, different catalog version
Memory isn't shown for this round. This machine's graphics chip shares the computer's memory in a way our measurement can't see, so the number would read far below what the model really uses.
General capability
| Model | Size | Tasks passed | Quality | Performance | Context |
|---|---|---|---|---|---|
| deepseek-v4-flash-284b-q2 | 284B | 33/33 (100%) | 8/10 (7 pieces) | 17.1 tok/s output 456 tok/s prefill | — |
| qwen3.8-27b-q4 (kv: f16) | 27B | 33/33 (100%) | 7.8/10 (9 pieces) | 31.1 tok/s output 708 tok/s prefill | 256K |
| ornith1.5-35b-a3b-q8 (kv: f16) | 35B | 29/33 (88%) | 7.7/10 (7 pieces) | 42.8 tok/s output 1,248 tok/s prefill | 64K |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context |
|---|---|---|---|---|---|
| deepseek-v4-flash-284b-q2 | 284B | 36/39 (92%) | 6.6/10 (27 pieces) | 17.1 tok/s output 456 tok/s prefill | — |
| ornith1.5-35b-a3b-q8 (kv: f16) | 35B | 36/39 (92%) | 7/10 (27 pieces) | 42.8 tok/s output 1,248 tok/s prefill | 64K |
| qwen3.8-27b-q4 (kv: f16) | 27B | 36/39 (92%) | 7.2/10 (27 pieces) | 31.1 tok/s output 708 tok/s prefill | 256K |
Engineering work
| Model | Size | Tasks passed | Quality | Performance | Context |
|---|---|---|---|---|---|
| ornith1.5-35b-a3b-q8 (kv: f16) | 35B | 19/30 (63%) | 6.1/10 (17 pieces) | 42.8 tok/s output 1,248 tok/s prefill | 64K |
| qwen3.8-27b-q4 (kv: f16) | 27B | 18/30 (60%) | 6.9/10 (15 pieces) | 31.1 tok/s output 708 tok/s prefill | 256K |
Not published — some tasks could not be measured on this round: deepseek-v4-flash-284b-q2 (2 task(s) unmeasured)
Complex workflows
| Model | Size | Tasks passed | Quality | Performance | Context |
|---|---|---|---|---|---|
| qwen3.8-27b-q4 (kv: f16) | 27B | 15/27 (56%) | 2.4/10 (6 pieces) | 29.1 tok/s output 713 tok/s prefill | 256K |
| ornith1.5-35b-a3b-q8 (kv: f16) | 35B | 12/26 (46%) | 2.4/10 (6 pieces) | 42.8 tok/s output 1,248 tok/s prefill | 64K |
Not published — some tasks could not be measured on this round: deepseek-v4-flash-284b-q2 (1 task(s) unmeasured)
Earlier round — 2026-08-30
DGX Spark Class · NVIDIA GB10 · 122 GB · llama-cpp · 3 trials per task · 2026-08-30 · gezel 8d7bb7e5 · catalog 0.1.49 — different gezel build, different catalog version
Memory isn't shown for this round. This machine's graphics chip shares the computer's memory in a way our measurement can't see, so the number would read far below what the model really uses.
General capability
| Model | Size | Tasks passed | Quality | Performance | Context |
|---|---|---|---|---|---|
| qwen3.8-27b-q4 (kv: f16) | 27B | 32/33 (97%) | 8.2/10 (9 pieces) | 32.4 tok/s output 777 tok/s prefill | 256K |
| qwen3.6-35b-a3b-q4 (kv: f16) | 35B | 31/33 (94%) | 8/10 (9 pieces) | 82 tok/s output 1,690 tok/s prefill | 64K |
| ornith1.5-9b-q8 (kv: f16) | 9B | 28/32 (88%) | 7.1/10 (9 pieces) | 25.6 tok/s output 2,117 tok/s prefill | 256K |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context |
|---|---|---|---|---|---|
| qwen3.8-27b-q4 (kv: f16) | 27B | 35/39 (90%) | 7.2/10 (27 pieces) | 32.4 tok/s output 777 tok/s prefill | 256K |
| qwen3.6-35b-a3b-q4 (kv: f16) | 35B | 34/39 (87%) | 6.9/10 (26 pieces) | 82 tok/s output 1,690 tok/s prefill | 64K |
| ornith1.5-9b-q8 (kv: f16) | 9B | 32/39 (82%) | 6.4/10 (27 pieces) | 25.6 tok/s output 2,117 tok/s prefill | 256K |
Earlier round — 2026-08-27
AMD Ryzen 9 7950X3D 16-Core Processor · AMD Radeon AI PRO R9700 · 64 GB · llama-cpp · 3 trials per task · 2026-08-27 · gezel 3cf0b7a2 · catalog 0.1.44 — different device, different gezel build, different catalog version
System memory counts this computer's own memory only. What the model keeps on the separate graphics card isn't included.
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | System memory |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q4 (kv: q8_0) | 27B | 32/33 (97%) | 8.1/10 (18 pieces) | 31.8 tok/s output 681 tok/s prefill | 256K | 15.8 GB |
| qwen3.6-35b-a3b-q4 (kv: q8_0) | 35B | 30/33 (91%) | 7.8/10 (9 pieces) | 131.1 tok/s output 2,102 tok/s prefill | 64K | 16.1 GB |
| gemma4-31b-q4 (kv: f16) | 30.7B | 29/33 (88%) | 7.1/10 (7 pieces) | 28.7 tok/s output 407 tok/s prefill | 99K | 16.8 GB |
| gemma4-26b-q4 (kv: f16) | 25.2B | 28/33 (85%) | 6.9/10 (6 pieces) | 131.2 tok/s output 1,820 tok/s prefill | 64K | 14.9 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | System memory |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q4 (kv: q8_0) | 27B | 36/39 (92%) | 7.5/10 (26 pieces) | 31.8 tok/s output 681 tok/s prefill | 256K | 14.9 GB |
| qwen3.6-35b-a3b-q4 (kv: q8_0) | 35B | 35/39 (90%) | 6.8/10 (26 pieces) | 131.1 tok/s output 2,102 tok/s prefill | 64K | 15.6 GB |
| gemma4-31b-q4 (kv: f16) | 30.7B | 30/39 (77%) | 6.5/10 (20 pieces) | 28.7 tok/s output 407 tok/s prefill | 99K | 16.9 GB |
| gemma4-26b-q4 (kv: f16) | 25.2B | 26/39 (67%) | 5.7/10 (22 pieces) | 131.2 tok/s output 1,820 tok/s prefill | 64K | 14.7 GB |
Earlier round — 2026-08-26
M4 Max · 64 GB · llama-cpp + mlx · 3 trials per task · 2026-08-26 · gezel d1d5e5c0 · catalog 0.1.42 — different device, different gezel build, different catalog version
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q2 (kv: q8_0) | 27B | 32/33 (97%) | 7.9/10 (9 pieces) | 24.1 tok/s output 241 tok/s prefill | 256K | 21.1 GB |
| qwen3.8-27b-q3 (kv: q8_0) | 27B | 32/33 (97%) | 8/10 (9 pieces) | 22.7 tok/s output 239 tok/s prefill | 256K | 24.5 GB |
| qwen3.8-27b-q4 (kv: q8_0) | 27B | 31/33 (94%) | 7.8/10 (17 pieces) | 21.9 tok/s output 232 tok/s prefill | 256K | 28.4 GB |
| qwen3.8-27b-q4 | 27B | 31/33 (94%) | 7.8/10 (17 pieces) | 68.4 tok/s output 232 tok/s prefill | — | — |
| qwen3.8-27b-iq1-s (kv: q8_0) | 27B | 14/33 (42%) | — | 11.1 tok/s output 190 tok/s prefill | 256K | 21.1 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q4 (kv: q8_0) | 27B | 36/39 (92%) | 7.3/10 (53 pieces) | 21.9 tok/s output 232 tok/s prefill | 256K | 26.2 GB |
| qwen3.8-27b-q4 | 27B | 36/39 (92%) | 7.3/10 (53 pieces) | 68.4 tok/s output 232 tok/s prefill | — | — |
| qwen3.8-27b-q2 (kv: q8_0) | 27B | 35/39 (90%) | 7.3/10 (25 pieces) | 24.1 tok/s output 241 tok/s prefill | 256K | 19.8 GB |
| qwen3.8-27b-q3 (kv: q8_0) | 27B | 34/39 (87%) | 7.1/10 (27 pieces) | 22.7 tok/s output 239 tok/s prefill | 256K | 24.4 GB |
| qwen3.8-27b-iq1-s (kv: q8_0) | 27B | 14/39 (36%) | 3.9/10 (25 pieces) | 11.1 tok/s output 190 tok/s prefill | 256K | 20.2 GB |
Earlier round — 2026-08-22
M4 Max · 64 GB · llama-cpp + mlx · 3 trials per task · 2026-08-22 · gezel b3be0a7f · catalog 0.1.39 — different device, different gezel build, different catalog version
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| gemma4-12b-q4 (kv: f16) | 12B | 31/33 (94%) | 6.9/10 (6 pieces) | 48.2 tok/s output 508 tok/s prefill | 99K | 42.2 GB |
| gemma4-12b-q8 (kv: f16) | 12B | 30/33 (91%) | 8.1/10 (3 pieces) | 30.3 tok/s output 506 tok/s prefill | 83K | 42.0 GB |
| btl4-compact-iq2 (kv: q8_0) | 35B | 29/33 (88%) | 7.1/10 (9 pieces) | 80.4 tok/s output 1,120 tok/s prefill | 256K | 14.6 GB |
| ornith1.5-9b-q8 (kv: q8_0) | 9B | 28/33 (85%) | 4.4/10 (5 pieces) | 47 tok/s output 823 tok/s prefill | 256K | 15.9 GB |
| qwen3.5-9b-q4 (kv: q8_0) | 9B | 28/33 (85%) | 5.3/10 (9 pieces) | 62.9 tok/s output 787 tok/s prefill | 256K | 13.1 GB |
| gemma4-e4b-q4 (kv: f16) | 8B | 26/33 (79%) | 5.4/10 (9 pieces) | 96.5 tok/s output 1,442 tok/s prefill | 128K | 13.7 GB |
| qwen3.5-4b-q4 (kv: q8_0) | 4B | 22/33 (67%) | 6.6/10 (9 pieces) | 89.3 tok/s output 1,336 tok/s prefill | 256K | 11.5 GB |
| gemma4-e2b-q4 (kv: f16) | 2.3B | 17/33 (52%) | 4.3/10 (9 pieces) | 144.1 tok/s output 2,574 tok/s prefill | 128K | 6.9 GB |
| lfm2.5-2.6b-q4 (kv: q8_0) | 2.6B | 13/33 (39%) | 3.6/10 (8 pieces) | 168.4 tok/s output 2,125 tok/s prefill | 125K | 4.9 GB |
| qwen3.5-2b-q4 (kv: q8_0) | 2B | 9/33 (27%) | 3.2/10 (9 pieces) | 155.3 tok/s output 3,084 tok/s prefill | 256K | 5.2 GB |
| mistral-7b-q4 (kv: q8_0) | 7B | 4/33 (12%) | — | 76 tok/s output 762 tok/s prefill | 32K | 9.6 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| gemma4-12b-q8 (kv: f16) | 12B | 35/39 (90%) | 5.8/10 (24 pieces) | 30.3 tok/s output 506 tok/s prefill | 83K | 41.8 GB |
| gemma4-12b-q4 (kv: f16) | 12B | 34/39 (87%) | 5.8/10 (24 pieces) | 48.2 tok/s output 508 tok/s prefill | 99K | 41.5 GB |
| ornith1.5-9b-q8 (kv: q8_0) | 9B | 34/39 (87%) | 6.5/10 (27 pieces) | 47 tok/s output 823 tok/s prefill | 256K | 14.9 GB |
| gemma4-e4b-q4 (kv: f16) | 8B | 31/39 (79%) | 5.5/10 (22 pieces) | 96.5 tok/s output 1,442 tok/s prefill | 128K | 12.5 GB |
| qwen3.5-9b-q4 (kv: q8_0) | 9B | 31/39 (79%) | 6.2/10 (25 pieces) | 62.9 tok/s output 787 tok/s prefill | 256K | 11.3 GB |
| qwen3.5-4b-q4 (kv: q8_0) | 4B | 30/39 (77%) | 6.2/10 (24 pieces) | 89.3 tok/s output 1,336 tok/s prefill | 256K | 9.9 GB |
| btl4-compact-iq2 (kv: q8_0) | 35B | 27/39 (69%) | 6.2/10 (25 pieces) | 80.4 tok/s output 1,120 tok/s prefill | 256K | 14.1 GB |
| gemma4-e2b-q4 (kv: f16) | 2.3B | 16/39 (41%) | 4.5/10 (20 pieces) | 144.1 tok/s output 2,574 tok/s prefill | 128K | 6.2 GB |
| lfm2.5-2.6b-q4 (kv: q8_0) | 2.6B | 13/39 (33%) | 4.4/10 (27 pieces) | 168.4 tok/s output 2,125 tok/s prefill | 125K | 4.8 GB |
| qwen3.5-2b-q4 (kv: q8_0) | 2B | 10/39 (26%) | 4.1/10 (26 pieces) | 155.3 tok/s output 3,084 tok/s prefill | 256K | 4.6 GB |
| mistral-7b-q4 (kv: q8_0) | 7B | 0/39 (0%) | 1.6/10 (30 pieces) | 76 tok/s output 762 tok/s prefill | 32K | 9.6 GB |
Earlier round — 2026-08-20
M4 Max · 64 GB · llama-cpp + mlx · 3 trials per task · 2026-08-20 · gezel 3bccbb5d · catalog 0.1.36+local.8c5b93111.qwen3.8-ud — different device, different gezel build, different catalog version
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q3 (kv: q8_0) | 27B | 33/33 (100%) | 8.5/10 (6 pieces) | 22.9 tok/s output 227 tok/s prefill | 256K | 24.1 GB |
| qwen3.8-27b-q4 (kv: q8_0) | 27B | 33/33 (100%) | 8.2/10 (9 pieces) | 22.9 tok/s output 227 tok/s prefill | 256K | 26.9 GB |
| gemma4-31b-q4 (kv: f16) | 30.7B | 32/33 (97%) | 6.2/10 (9 pieces) | 22.1 tok/s output 196 tok/s prefill | 256K | 45.8 GB |
| qwen3.8-27b-q6 (kv: q8_0) | 27B | 32/33 (97%) | 8/10 (7 pieces) | 18.8 tok/s output 227 tok/s prefill | 256K | 33.2 GB |
| qwen3.8-27b-q8 (kv: q8_0) | 27B | 32/33 (97%) | 8.3/10 (6 pieces) | 13.4 tok/s output 219 tok/s prefill | 256K | 38.5 GB |
| qwen3.6-35b-a3b-q4 (kv: q8_0) | 35B | 31/33 (94%) | 8/10 (9 pieces) | 75.6 tok/s output 1,156 tok/s prefill | 256K | 25.4 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q3 (kv: q8_0) | 27B | 35/39 (90%) | 7.5/10 (24 pieces) | 22.9 tok/s output 227 tok/s prefill | 256K | 22.9 GB |
| qwen3.8-27b-q4 (kv: q8_0) | 27B | 35/39 (90%) | 6.9/10 (25 pieces) | 22.9 tok/s output 227 tok/s prefill | 256K | 27.1 GB |
| qwen3.8-27b-q6 (kv: q8_0) | 27B | 35/39 (90%) | 7.1/10 (25 pieces) | 18.8 tok/s output 227 tok/s prefill | 256K | 32.6 GB |
| qwen3.6-35b-a3b-q4 (kv: q8_0) | 35B | 33/39 (85%) | 6.9/10 (24 pieces) | 75.6 tok/s output 1,156 tok/s prefill | 256K | 25.1 GB |
| qwen3.8-27b-q8 (kv: q8_0) | 27B | 33/39 (85%) | 7.3/10 (24 pieces) | 13.4 tok/s output 219 tok/s prefill | 256K | 37.6 GB |
| gemma4-31b-q4 (kv: f16) | 30.7B | 28/39 (72%) | 6.4/10 (19 pieces) | 22.1 tok/s output 196 tok/s prefill | 256K | 45.7 GB |
Earlier round — 2026-08-20
DGX Spark Class · 122 GB · ds4 + llama-cpp · 3 trials per task · 2026-08-20 · gezel 2745e97 · catalog 0.1.36 — different device, different gezel build, different catalog version
Memory isn't shown for this round. This machine's graphics chip shares the computer's memory in a way our measurement can't see, so the number would read far below what the model really uses.
General capability
| Model | Size | Tasks passed | Quality | Performance | Context |
|---|---|---|---|---|---|
| deepseek-v4-flash-284b-q2 | 284B | 33/33 (100%) | 7.8/10 (7 pieces) | 16.9 tok/s output 680 tok/s prefill | — |
| qwen3.8-27b-q4 (kv: q8_0) | 27B | 32/32 (100%) | 8.2/10 (9 pieces) | 12.1 tok/s output 706 tok/s prefill | 256K |
| qwen3.6-27b-q8 (kv: q8_0) | 27B | 30/31 (97%) | 8.2/10 (9 pieces) | 7.8 tok/s output 725 tok/s prefill | 256K |
| btl4-35b-q4 (kv: q8_0) | 35B | 27/31 (87%) | 7.8/10 (8 pieces) | 73.2 tok/s output 1,608 tok/s prefill | 64K |
| gemma4-e4b-q4 (kv: f16) | 8B | 25/33 (76%) | 6.2/10 (3 pieces) | 65.4 tok/s output 4,698 tok/s prefill | 128K |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context |
|---|---|---|---|---|---|
| qwen3.8-27b-q4 (kv: q8_0) | 27B | 35/39 (90%) | 7.2/10 (27 pieces) | 12.1 tok/s output 706 tok/s prefill | 256K |
| deepseek-v4-flash-284b-q2 | 284B | 34/39 (87%) | 6.5/10 (26 pieces) | 16.9 tok/s output 680 tok/s prefill | — |
| qwen3.6-27b-q8 (kv: q8_0) | 27B | 32/39 (82%) | 6.8/10 (27 pieces) | 7.8 tok/s output 725 tok/s prefill | 256K |
| btl4-35b-q4 (kv: q8_0) | 35B | 31/39 (79%) | 6.7/10 (27 pieces) | 73.2 tok/s output 1,608 tok/s prefill | 64K |
| gemma4-e4b-q4 (kv: f16) | 8B | 31/39 (79%) | 5.2/10 (26 pieces) | 65.4 tok/s output 4,698 tok/s prefill | 128K |
Earlier round — 2026-08-14
M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-14 · gezel e367f442 · catalog 0.1.29 — different device, different gezel build, different catalog version
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q4 | 27B | 33/33 (100%) | 8.2/10 (9 pieces) | 22.4 tok/s output 233 tok/s prefill | 256K | 27.6 GB |
| qwen3.6-27b-q4 | 27B | 32/33 (97%) | 7.9/10 (9 pieces) | 22.9 tok/s output 230 tok/s prefill | 256K | 29.0 GB |
| qwen3.6-35b-a3b-q4 | 35B | 31/33 (94%) | 8/10 (9 pieces) | 78.5 tok/s output 1,182 tok/s prefill | 256K | 26.0 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q4 | 27B | 35/39 (90%) | 7.1/10 (26 pieces) | 22.4 tok/s output 233 tok/s prefill | 256K | 27.3 GB |
| qwen3.6-35b-a3b-q4 | 35B | 34/39 (87%) | 7/10 (27 pieces) | 78.5 tok/s output 1,182 tok/s prefill | 256K | 25.5 GB |
Not published — some tasks could not be measured on this round: qwen3.6-27b-q4 (1 task(s) unmeasured)
Earlier round — 2026-08-13
M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-13 · gezel c5904085 · catalog 0.1.23 — different device, different gezel build, different catalog version
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| gemma4-26b-q4 | 25.2B | 33/33 (100%) | 5.7/10 (9 pieces) | 107.5 tok/s output 1,264 tok/s prefill | 118K | 41.7 GB |
| muse-glimmer-30b-q4 | 30B | 30/33 (91%) | 6.5/10 (9 pieces) | 26 tok/s output 226 tok/s prefill | 128K | 19.2 GB |
| qwen3.6-35b-a3b-q4 | 35B | 30/33 (91%) | 8.1/10 (9 pieces) | 71.3 tok/s output 1,104 tok/s prefill | 256K | 25.9 GB |
| nemotron3.5-lightning-30b-q4 | 30B | 26/33 (79%) | 5.6/10 (9 pieces) | 93.2 tok/s output 1,030 tok/s prefill | 1024K | 29.4 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.6-35b-a3b-q4 | 35B | 33/39 (85%) | 6.6/10 (27 pieces) | 71.3 tok/s output 1,104 tok/s prefill | 256K | 25.1 GB |
| gemma4-26b-q4 | 25.2B | 28/39 (72%) | 6/10 (19 pieces) | 107.5 tok/s output 1,264 tok/s prefill | 118K | 40.8 GB |
| muse-glimmer-30b-q4 | 30B | 25/39 (64%) | 5.5/10 (25 pieces) | 26 tok/s output 228 tok/s prefill | 128K | 19.2 GB |
| nemotron3.5-lightning-30b-q4 | 30B | 23/39 (59%) | 5.6/10 (27 pieces) | 93.2 tok/s output 1,030 tok/s prefill | 1024K | 28.8 GB |
Earlier round — 2026-08-11
M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-11 · gezel b2ea2819 · catalog 0.1.20 — different device, different gezel build, different catalog version
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| gemma4-26b-q4 | 25.2B | 33/33 (100%) | 6.4/10 (9 pieces) | 89.6 tok/s output 1,093 tok/s prefill | 118K | 40.9 GB |
| qwen3.6-27b-q4 | 27B | 33/33 (100%) | 8/10 (9 pieces) | 18.5 tok/s output 198 tok/s prefill | 256K | 26.4 GB |
| qwen3.6-35b-a3b-q4 | 35B | 32/33 (97%) | 8.2/10 (9 pieces) | 76.4 tok/s output 1,125 tok/s prefill | 256K | 26.5 GB |
| gemma4-31b-q4 | 30.7B | 31/33 (94%) | 6.7/10 (9 pieces) | 23.3 tok/s output 202 tok/s prefill | 256K | 43.1 GB |
| gemma4-e4b-q4 | 8B | 31/33 (94%) | 4.7/10 (9 pieces) | 94.7 tok/s output 1,398 tok/s prefill | 128K | 13.5 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.6-35b-a3b-q4 | 35B | 33/39 (85%) | 6.6/10 (26 pieces) | 76.4 tok/s output 1,125 tok/s prefill | 256K | 25.6 GB |
| gemma4-26b-q4 | 25.2B | 30/39 (77%) | 6.1/10 (20 pieces) | 89.6 tok/s output 1,093 tok/s prefill | 118K | 40.2 GB |
| gemma4-e4b-q4 | 8B | 29/39 (74%) | 5.4/10 (25 pieces) | 94.7 tok/s output 1,398 tok/s prefill | 128K | 12.4 GB |
| qwen3.6-27b-q4 | 27B | 29/39 (74%) | 5.6/10 (24 pieces) | 18.5 tok/s output 198 tok/s prefill | 256K | 27.0 GB |
| gemma4-31b-q4 | 30.7B | 28/39 (72%) | 6.1/10 (20 pieces) | 23.3 tok/s output 202 tok/s prefill | 256K | 45.1 GB |
Earlier round — 2026-08-09
M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-09 · gezel e2859602 · catalog 0.1.17 — different device, different gezel build, different catalog version
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| gemma4-31b-q4 | 30.7B | 33/33 (100%) | 6.8/10 (9 pieces) | 22.3 tok/s output 195 tok/s prefill | 64K | 31.0 GB |
| gemma4-12b-q4 | 12B | 28/33 (85%) | 6.9/10 (9 pieces) | 50.9 tok/s output 543 tok/s prefill | 64K | 30.6 GB |
| gemma4-e4b-q4 | 8B | 24/33 (73%) | 4.6/10 (9 pieces) | 97.8 tok/s output 1,411 tok/s prefill | 64K | 10.4 GB |
| ornith-9b-q4 | 9B | 24/33 (73%) | 6.7/10 (9 pieces) | 64.9 tok/s output 657 tok/s prefill | 64K | 9.8 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| ornith-9b-q4 | 9B | 34/39 (87%) | 6.5/10 (26 pieces) | 64.9 tok/s output 657 tok/s prefill | 64K | 8.9 GB |
| gemma4-12b-q4 | 12B | 29/39 (74%) | 4.8/10 (26 pieces) | 50.9 tok/s output 543 tok/s prefill | 64K | 29.7 GB |
| gemma4-e4b-q4 | 8B | 28/39 (72%) | 5.2/10 (22 pieces) | 97.8 tok/s output 1,411 tok/s prefill | 64K | 8.9 GB |
| gemma4-31b-q4 | 30.7B | 25/39 (64%) | 6.5/10 (18 pieces) | 22.3 tok/s output 195 tok/s prefill | 64K | 30.1 GB |
Reading the table
Tasks passed counts every attempt across every job in the set. A model with 24/33 (73%) finished 24 of 33 attempts correctly. If any job in the set ran fewer than three times you'll see a raw count instead of a percentage — too small a sample to quote as a rate. A model whose results were incomplete is left out of the table entirely rather than shown with a gap.
Quality is an AI reviewer's opinion of the finished work, and the count beside it is how many pieces that opinion covers. It only grades work that was actually produced, so a model that fails often is judged on its successes alone — 6.5/10 (18 pieces) is a weaker claim than 6.5/10 (26 pieces). Treat it as colour next to the pass rate, never as a substitute for it.
Performance shows two measured speeds on the machine named above: output speed is how fast the model writes its answer, and prefill speed is how fast it takes in your prompt and documents. Both matter for how a gezel feels — prefill governs the pause before it starts, while output governs how fast text appears. A dash means that round did not record a throughput probe.
Context is the working memory the model was given for these runs — how much it can hold at once. Memory used is the most memory the model and its engine held on the test machine. What that number covers depends on the machine, so each round says which one it measured:
On a Mac the graphics chip shares the computer's memory and our measurement sees all of it, so Memory used is the model's whole footprint.
On a PC with a separate graphics card the column is called System memory. It counts the computer's own memory only, not what the model keeps on the graphics card.
On machines like the DGX Spark, where the graphics chip shares memory in a way our measurement can't see, the column is left out. The number would read far below what the model really uses.
Treat these as a rough guide from one machine, not a promise about yours. The Models catalogue in the app is the better check for whether a model fits your computer. When the engine reported its KV-cache precision, it appears beside the model name — for example, (kv: q8_0).
Each test round keeps all of its tables together under one provenance stamp. Earlier rounds stay separate from the latest because a change to gezel or to the task set can move a score without any model changing.
Size is the model's parameter count where we know it. Bigger is often but not always better: on office work in particular, some smaller models beat larger ones, and a model family's habits matter more than its size.
Choosing from this
A high score on the office set is the better guide for everyday document, planning, and analysis work. A high score on the general set matters more if you want a gezel writing or fixing code, and the engineering set, where a round has it, is the stricter test of that.
If a model you're considering isn't listed, it hasn't been measured here yet — which is not a verdict on it either way. The Models catalogue in the app will still tell you whether it fits this machine.