Model scorecard
These are measured results, not estimates. Each number is the count of jobs a model actually finished correctly, checked by a program on a real machine.
How we test models explains what was counted and what these numbers do not tell you. The short version: each job is run three times, a job passes only when every requirement is met, and runs lost to machine trouble are set aside rather than blamed on the model.
Results by test round
Each round has two tables measured on the same machine and build. General capability covers writing a small working program, fixing a bug from its symptoms, following a procedure and stopping at a problem, and turning several documents into one reconciled summary. Office and knowledge work covers a customer notice under a hard word limit, a meeting turned into an action register, a cited research brief, an experiment read-out, a spreadsheet model, a slide deck, and a Word document.
Latest round — 2026-08-30
DGX Spark Class · NVIDIA GB10 · 122 GB · llama-cpp · 3 trials per task · 2026-08-30 · gezel 8d7bb7e5 · catalog 0.1.49
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q4 (kv: f16) | 27B | 32/33 (97%) | 8.2/10 (9 pieces) | 32.4 tok/s output 777 tok/s prefill | 256K | 6.1 GB |
| qwen3.6-35b-a3b-q4 (kv: f16) | 35B | 31/33 (94%) | 8/10 (9 pieces) | 82 tok/s output 1,690 tok/s prefill | 64K | 4.2 GB |
| ornith1.5-9b-q8 (kv: f16) | 9B | 28/32 (88%) | 7.1/10 (9 pieces) | 25.6 tok/s output 2,117 tok/s prefill | 256K | 5.2 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q4 (kv: f16) | 27B | 35/39 (90%) | 7.2/10 (27 pieces) | 32.4 tok/s output 777 tok/s prefill | 256K | 4.9 GB |
| qwen3.6-35b-a3b-q4 (kv: f16) | 35B | 34/39 (87%) | 6.9/10 (26 pieces) | 82 tok/s output 1,690 tok/s prefill | 64K | 4.1 GB |
| ornith1.5-9b-q8 (kv: f16) | 9B | 32/39 (82%) | 6.4/10 (27 pieces) | 25.6 tok/s output 2,117 tok/s prefill | 256K | 4.5 GB |
Earlier round — 2026-08-27
AMD Ryzen 9 7950X3D 16-Core Processor · AMD Radeon AI PRO R9700 · 64 GB · llama-cpp · 3 trials per task · 2026-08-27 · gezel 3cf0b7a2 · catalog 0.1.44 — different device, different gezel build, different catalog version
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q4 (kv: q8_0) | 27B | 32/33 (97%) | 8.1/10 (18 pieces) | 31.8 tok/s output 681 tok/s prefill | 256K | 15.8 GB |
| qwen3.6-35b-a3b-q4 (kv: q8_0) | 35B | 30/33 (91%) | 7.8/10 (9 pieces) | 131.1 tok/s output 2,102 tok/s prefill | 64K | 16.1 GB |
| gemma4-31b-q4 (kv: f16) | 30.7B | 29/33 (88%) | 7.1/10 (7 pieces) | 28.7 tok/s output 407 tok/s prefill | 99K | 16.8 GB |
| gemma4-26b-q4 (kv: f16) | 25.2B | 28/33 (85%) | 6.9/10 (6 pieces) | 131.2 tok/s output 1,820 tok/s prefill | 64K | 14.9 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q4 (kv: q8_0) | 27B | 36/39 (92%) | 7.5/10 (26 pieces) | 31.8 tok/s output 681 tok/s prefill | 256K | 14.9 GB |
| qwen3.6-35b-a3b-q4 (kv: q8_0) | 35B | 35/39 (90%) | 6.8/10 (26 pieces) | 131.1 tok/s output 2,102 tok/s prefill | 64K | 15.6 GB |
| gemma4-31b-q4 (kv: f16) | 30.7B | 30/39 (77%) | 6.5/10 (20 pieces) | 28.7 tok/s output 407 tok/s prefill | 99K | 16.9 GB |
| gemma4-26b-q4 (kv: f16) | 25.2B | 26/39 (67%) | 5.7/10 (22 pieces) | 131.2 tok/s output 1,820 tok/s prefill | 64K | 14.7 GB |
Earlier round — 2026-08-26
M4 Max · 64 GB · llama-cpp + mlx · 3 trials per task · 2026-08-26 · gezel d1d5e5c0 · catalog 0.1.42 — different device, different gezel build, different catalog version
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q2 (kv: q8_0) | 27B | 32/33 (97%) | 7.9/10 (9 pieces) | 24.1 tok/s output 241 tok/s prefill | 256K | 21.1 GB |
| qwen3.8-27b-q3 (kv: q8_0) | 27B | 32/33 (97%) | 8/10 (9 pieces) | 22.7 tok/s output 239 tok/s prefill | 256K | 24.5 GB |
| qwen3.8-27b-q4 (kv: q8_0) | 27B | 31/33 (94%) | 7.8/10 (17 pieces) | 21.9 tok/s output 232 tok/s prefill | 256K | 28.4 GB |
| qwen3.8-27b-q4 | 27B | 31/33 (94%) | 7.8/10 (17 pieces) | 68.4 tok/s output 232 tok/s prefill | — | — |
| qwen3.8-27b-iq1-s (kv: q8_0) | 27B | 14/33 (42%) | — | 11.1 tok/s output 190 tok/s prefill | 256K | 21.1 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q4 (kv: q8_0) | 27B | 36/39 (92%) | 7.3/10 (53 pieces) | 21.9 tok/s output 232 tok/s prefill | 256K | 26.2 GB |
| qwen3.8-27b-q4 | 27B | 36/39 (92%) | 7.3/10 (53 pieces) | 68.4 tok/s output 232 tok/s prefill | — | — |
| qwen3.8-27b-q2 (kv: q8_0) | 27B | 35/39 (90%) | 7.3/10 (25 pieces) | 24.1 tok/s output 241 tok/s prefill | 256K | 19.8 GB |
| qwen3.8-27b-q3 (kv: q8_0) | 27B | 34/39 (87%) | 7.1/10 (27 pieces) | 22.7 tok/s output 239 tok/s prefill | 256K | 24.4 GB |
| qwen3.8-27b-iq1-s (kv: q8_0) | 27B | 14/39 (36%) | 3.9/10 (25 pieces) | 11.1 tok/s output 190 tok/s prefill | 256K | 20.2 GB |
Earlier round — 2026-08-22
M4 Max · 64 GB · llama-cpp + mlx · 3 trials per task · 2026-08-22 · gezel b3be0a7f · catalog 0.1.39 — different device, different gezel build, different catalog version
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| gemma4-12b-q4 (kv: f16) | 12B | 31/33 (94%) | 6.9/10 (6 pieces) | 48.2 tok/s output 508 tok/s prefill | 99K | 42.2 GB |
| gemma4-12b-q8 (kv: f16) | 12B | 30/33 (91%) | 8.1/10 (3 pieces) | 30.3 tok/s output 506 tok/s prefill | 83K | 42.0 GB |
| btl4-compact-iq2 (kv: q8_0) | 35B | 29/33 (88%) | 7.1/10 (9 pieces) | 80.4 tok/s output 1,120 tok/s prefill | 256K | 14.6 GB |
| ornith1.5-9b-q8 (kv: q8_0) | 9B | 28/33 (85%) | 4.4/10 (5 pieces) | 47 tok/s output 823 tok/s prefill | 256K | 15.9 GB |
| qwen3.5-9b-q4 (kv: q8_0) | 9B | 28/33 (85%) | 5.3/10 (9 pieces) | 62.9 tok/s output 787 tok/s prefill | 256K | 13.1 GB |
| gemma4-e4b-q4 (kv: f16) | 8B | 26/33 (79%) | 5.4/10 (9 pieces) | 96.5 tok/s output 1,442 tok/s prefill | 128K | 13.7 GB |
| qwen3.5-4b-q4 (kv: q8_0) | 4B | 22/33 (67%) | 6.6/10 (9 pieces) | 89.3 tok/s output 1,336 tok/s prefill | 256K | 11.5 GB |
| gemma4-e2b-q4 (kv: f16) | 2.3B | 17/33 (52%) | 4.3/10 (9 pieces) | 144.1 tok/s output 2,574 tok/s prefill | 128K | 6.9 GB |
| lfm2.5-2.6b-q4 (kv: q8_0) | 2.6B | 13/33 (39%) | 3.6/10 (8 pieces) | 168.4 tok/s output 2,125 tok/s prefill | 125K | 4.9 GB |
| qwen3.5-2b-q4 (kv: q8_0) | 2B | 9/33 (27%) | 3.2/10 (9 pieces) | 155.3 tok/s output 3,084 tok/s prefill | 256K | 5.2 GB |
| mistral-7b-q4 (kv: q8_0) | 7B | 4/33 (12%) | — | 76 tok/s output 762 tok/s prefill | 32K | 9.6 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| gemma4-12b-q8 (kv: f16) | 12B | 35/39 (90%) | 5.8/10 (24 pieces) | 30.3 tok/s output 506 tok/s prefill | 83K | 41.8 GB |
| gemma4-12b-q4 (kv: f16) | 12B | 34/39 (87%) | 5.8/10 (24 pieces) | 48.2 tok/s output 508 tok/s prefill | 99K | 41.5 GB |
| ornith1.5-9b-q8 (kv: q8_0) | 9B | 34/39 (87%) | 6.5/10 (27 pieces) | 47 tok/s output 823 tok/s prefill | 256K | 14.9 GB |
| gemma4-e4b-q4 (kv: f16) | 8B | 31/39 (79%) | 5.5/10 (22 pieces) | 96.5 tok/s output 1,442 tok/s prefill | 128K | 12.5 GB |
| qwen3.5-9b-q4 (kv: q8_0) | 9B | 31/39 (79%) | 6.2/10 (25 pieces) | 62.9 tok/s output 787 tok/s prefill | 256K | 11.3 GB |
| qwen3.5-4b-q4 (kv: q8_0) | 4B | 30/39 (77%) | 6.2/10 (24 pieces) | 89.3 tok/s output 1,336 tok/s prefill | 256K | 9.9 GB |
| btl4-compact-iq2 (kv: q8_0) | 35B | 27/39 (69%) | 6.2/10 (25 pieces) | 80.4 tok/s output 1,120 tok/s prefill | 256K | 14.1 GB |
| gemma4-e2b-q4 (kv: f16) | 2.3B | 16/39 (41%) | 4.5/10 (20 pieces) | 144.1 tok/s output 2,574 tok/s prefill | 128K | 6.2 GB |
| lfm2.5-2.6b-q4 (kv: q8_0) | 2.6B | 13/39 (33%) | 4.4/10 (27 pieces) | 168.4 tok/s output 2,125 tok/s prefill | 125K | 4.8 GB |
| qwen3.5-2b-q4 (kv: q8_0) | 2B | 10/39 (26%) | 4.1/10 (26 pieces) | 155.3 tok/s output 3,084 tok/s prefill | 256K | 4.6 GB |
| mistral-7b-q4 (kv: q8_0) | 7B | 0/39 (0%) | 1.6/10 (30 pieces) | 76 tok/s output 762 tok/s prefill | 32K | 9.6 GB |
Earlier round — 2026-08-20
M4 Max · 64 GB · llama-cpp + mlx · 3 trials per task · 2026-08-20 · gezel 3bccbb5d · catalog 0.1.36+local.8c5b93111.qwen3.8-ud — different device, different gezel build, different catalog version
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q3 (kv: q8_0) | 27B | 33/33 (100%) | 8.5/10 (6 pieces) | 22.9 tok/s output 227 tok/s prefill | 256K | 24.1 GB |
| qwen3.8-27b-q4 (kv: q8_0) | 27B | 33/33 (100%) | 8.2/10 (9 pieces) | 22.9 tok/s output 227 tok/s prefill | 256K | 26.9 GB |
| gemma4-31b-q4 (kv: f16) | 30.7B | 32/33 (97%) | 6.2/10 (9 pieces) | 22.1 tok/s output 196 tok/s prefill | 256K | 45.8 GB |
| qwen3.8-27b-q6 (kv: q8_0) | 27B | 32/33 (97%) | 8/10 (7 pieces) | 18.8 tok/s output 227 tok/s prefill | 256K | 33.2 GB |
| qwen3.8-27b-q8 (kv: q8_0) | 27B | 32/33 (97%) | 8.3/10 (6 pieces) | 13.4 tok/s output 219 tok/s prefill | 256K | 38.5 GB |
| qwen3.6-35b-a3b-q4 (kv: q8_0) | 35B | 31/33 (94%) | 8/10 (9 pieces) | 75.6 tok/s output 1,156 tok/s prefill | 256K | 25.4 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q3 (kv: q8_0) | 27B | 35/39 (90%) | 7.5/10 (24 pieces) | 22.9 tok/s output 227 tok/s prefill | 256K | 22.9 GB |
| qwen3.8-27b-q4 (kv: q8_0) | 27B | 35/39 (90%) | 6.9/10 (25 pieces) | 22.9 tok/s output 227 tok/s prefill | 256K | 27.1 GB |
| qwen3.8-27b-q6 (kv: q8_0) | 27B | 35/39 (90%) | 7.1/10 (25 pieces) | 18.8 tok/s output 227 tok/s prefill | 256K | 32.6 GB |
| qwen3.6-35b-a3b-q4 (kv: q8_0) | 35B | 33/39 (85%) | 6.9/10 (24 pieces) | 75.6 tok/s output 1,156 tok/s prefill | 256K | 25.1 GB |
| qwen3.8-27b-q8 (kv: q8_0) | 27B | 33/39 (85%) | 7.3/10 (24 pieces) | 13.4 tok/s output 219 tok/s prefill | 256K | 37.6 GB |
| gemma4-31b-q4 (kv: f16) | 30.7B | 28/39 (72%) | 6.4/10 (19 pieces) | 22.1 tok/s output 196 tok/s prefill | 256K | 45.7 GB |
Earlier round — 2026-08-20
DGX Spark Class · 122 GB · ds4 + llama-cpp · 3 trials per task · 2026-08-20 · gezel 2745e97 · catalog 0.1.36 — different device, different gezel build, different catalog version
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| deepseek-v4-flash-284b-q2 | 284B | 33/33 (100%) | 7.8/10 (7 pieces) | 16.9 tok/s output 680 tok/s prefill | — | — |
| qwen3.8-27b-q4 (kv: q8_0) | 27B | 32/32 (100%) | 8.2/10 (9 pieces) | 12.1 tok/s output 706 tok/s prefill | 256K | 18.0 GB |
| qwen3.6-27b-q8 (kv: q8_0) | 27B | 30/31 (97%) | 8.2/10 (9 pieces) | 7.8 tok/s output 725 tok/s prefill | 256K | 29.3 GB |
| btl4-35b-q4 (kv: q8_0) | 35B | 27/31 (87%) | 7.8/10 (8 pieces) | 73.2 tok/s output 1,608 tok/s prefill | 64K | 22.2 GB |
| gemma4-e4b-q4 (kv: f16) | 8B | 25/33 (76%) | 6.2/10 (3 pieces) | 65.4 tok/s output 4,698 tok/s prefill | 128K | 5.7 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q4 (kv: q8_0) | 27B | 35/39 (90%) | 7.2/10 (27 pieces) | 12.1 tok/s output 706 tok/s prefill | 256K | 5.4 GB |
| deepseek-v4-flash-284b-q2 | 284B | 34/39 (87%) | 6.5/10 (26 pieces) | 16.9 tok/s output 680 tok/s prefill | — | — |
| qwen3.6-27b-q8 (kv: q8_0) | 27B | 32/39 (82%) | 6.8/10 (27 pieces) | 7.8 tok/s output 725 tok/s prefill | 256K | 29.2 GB |
| btl4-35b-q4 (kv: q8_0) | 35B | 31/39 (79%) | 6.7/10 (27 pieces) | 73.2 tok/s output 1,608 tok/s prefill | 64K | 4.8 GB |
| gemma4-e4b-q4 (kv: f16) | 8B | 31/39 (79%) | 5.2/10 (26 pieces) | 65.4 tok/s output 4,698 tok/s prefill | 128K | 5.1 GB |
Earlier round — 2026-08-14
M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-14 · gezel e367f442 · catalog 0.1.29 — different device, different gezel build, different catalog version
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q4 | 27B | 33/33 (100%) | 8.2/10 (9 pieces) | 22.4 tok/s output 233 tok/s prefill | 256K | 27.6 GB |
| qwen3.6-27b-q4 | 27B | 32/33 (97%) | 7.9/10 (9 pieces) | 22.9 tok/s output 230 tok/s prefill | 256K | 29.0 GB |
| qwen3.6-35b-a3b-q4 | 35B | 31/33 (94%) | 8/10 (9 pieces) | 78.5 tok/s output 1,182 tok/s prefill | 256K | 26.0 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.8-27b-q4 | 27B | 35/39 (90%) | 7.1/10 (26 pieces) | 22.4 tok/s output 233 tok/s prefill | 256K | 27.3 GB |
| qwen3.6-35b-a3b-q4 | 35B | 34/39 (87%) | 7/10 (27 pieces) | 78.5 tok/s output 1,182 tok/s prefill | 256K | 25.5 GB |
Not published — some tasks could not be measured on this round: qwen3.6-27b-q4 (1 task(s) unmeasured)
Earlier round — 2026-08-13
M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-13 · gezel c5904085 · catalog 0.1.23 — different device, different gezel build, different catalog version
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| gemma4-26b-q4 | 25.2B | 33/33 (100%) | 5.7/10 (9 pieces) | 107.5 tok/s output 1,264 tok/s prefill | 118K | 41.7 GB |
| muse-glimmer-30b-q4 | 30B | 30/33 (91%) | 6.5/10 (9 pieces) | 26 tok/s output 226 tok/s prefill | 128K | 19.2 GB |
| qwen3.6-35b-a3b-q4 | 35B | 30/33 (91%) | 8.1/10 (9 pieces) | 71.3 tok/s output 1,104 tok/s prefill | 256K | 25.9 GB |
| nemotron3.5-lightning-30b-q4 | 30B | 26/33 (79%) | 5.6/10 (9 pieces) | 93.2 tok/s output 1,030 tok/s prefill | 1024K | 29.4 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.6-35b-a3b-q4 | 35B | 33/39 (85%) | 6.6/10 (27 pieces) | 71.3 tok/s output 1,104 tok/s prefill | 256K | 25.1 GB |
| gemma4-26b-q4 | 25.2B | 28/39 (72%) | 6/10 (19 pieces) | 107.5 tok/s output 1,264 tok/s prefill | 118K | 40.8 GB |
| muse-glimmer-30b-q4 | 30B | 25/39 (64%) | 5.5/10 (25 pieces) | 26 tok/s output 228 tok/s prefill | 128K | 19.2 GB |
| nemotron3.5-lightning-30b-q4 | 30B | 23/39 (59%) | 5.6/10 (27 pieces) | 93.2 tok/s output 1,030 tok/s prefill | 1024K | 28.8 GB |
Earlier round — 2026-08-11
M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-11 · gezel b2ea2819 · catalog 0.1.20 — different device, different gezel build, different catalog version
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| gemma4-26b-q4 | 25.2B | 33/33 (100%) | 6.4/10 (9 pieces) | 89.6 tok/s output 1,093 tok/s prefill | 118K | 40.9 GB |
| qwen3.6-27b-q4 | 27B | 33/33 (100%) | 8/10 (9 pieces) | 18.5 tok/s output 198 tok/s prefill | 256K | 26.4 GB |
| qwen3.6-35b-a3b-q4 | 35B | 32/33 (97%) | 8.2/10 (9 pieces) | 76.4 tok/s output 1,125 tok/s prefill | 256K | 26.5 GB |
| gemma4-31b-q4 | 30.7B | 31/33 (94%) | 6.7/10 (9 pieces) | 23.3 tok/s output 202 tok/s prefill | 256K | 43.1 GB |
| gemma4-e4b-q4 | 8B | 31/33 (94%) | 4.7/10 (9 pieces) | 94.7 tok/s output 1,398 tok/s prefill | 128K | 13.5 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| qwen3.6-35b-a3b-q4 | 35B | 33/39 (85%) | 6.6/10 (26 pieces) | 76.4 tok/s output 1,125 tok/s prefill | 256K | 25.6 GB |
| gemma4-26b-q4 | 25.2B | 30/39 (77%) | 6.1/10 (20 pieces) | 89.6 tok/s output 1,093 tok/s prefill | 118K | 40.2 GB |
| gemma4-e4b-q4 | 8B | 29/39 (74%) | 5.4/10 (25 pieces) | 94.7 tok/s output 1,398 tok/s prefill | 128K | 12.4 GB |
| qwen3.6-27b-q4 | 27B | 29/39 (74%) | 5.6/10 (24 pieces) | 18.5 tok/s output 198 tok/s prefill | 256K | 27.0 GB |
| gemma4-31b-q4 | 30.7B | 28/39 (72%) | 6.1/10 (20 pieces) | 23.3 tok/s output 202 tok/s prefill | 256K | 45.1 GB |
Earlier round — 2026-08-09
M4 Max · 64 GB · llama-cpp · 3 trials per task · 2026-08-09 · gezel e2859602 · catalog 0.1.17 — different device, different gezel build, different catalog version
General capability
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| gemma4-31b-q4 | 30.7B | 33/33 (100%) | 6.8/10 (9 pieces) | 22.3 tok/s output 195 tok/s prefill | 64K | 31.0 GB |
| gemma4-12b-q4 | 12B | 28/33 (85%) | 6.9/10 (9 pieces) | 50.9 tok/s output 543 tok/s prefill | 64K | 30.6 GB |
| gemma4-e4b-q4 | 8B | 24/33 (73%) | 4.6/10 (9 pieces) | 97.8 tok/s output 1,411 tok/s prefill | 64K | 10.4 GB |
| ornith-9b-q4 | 9B | 24/33 (73%) | 6.7/10 (9 pieces) | 64.9 tok/s output 657 tok/s prefill | 64K | 9.8 GB |
Office and knowledge work
| Model | Size | Tasks passed | Quality | Performance | Context | Memory used |
|---|---|---|---|---|---|---|
| ornith-9b-q4 | 9B | 34/39 (87%) | 6.5/10 (26 pieces) | 64.9 tok/s output 657 tok/s prefill | 64K | 8.9 GB |
| gemma4-12b-q4 | 12B | 29/39 (74%) | 4.8/10 (26 pieces) | 50.9 tok/s output 543 tok/s prefill | 64K | 29.7 GB |
| gemma4-e4b-q4 | 8B | 28/39 (72%) | 5.2/10 (22 pieces) | 97.8 tok/s output 1,411 tok/s prefill | 64K | 8.9 GB |
| gemma4-31b-q4 | 30.7B | 25/39 (64%) | 6.5/10 (18 pieces) | 22.3 tok/s output 195 tok/s prefill | 64K | 30.1 GB |
Reading the table
Tasks passed counts every attempt across every job in the set. A model with 24/33 (73%) finished 24 of 33 attempts correctly. If any job in the set ran fewer than three times you'll see a raw count instead of a percentage — too small a sample to quote as a rate. A model whose results were incomplete is left out of the table entirely rather than shown with a gap.
Quality is an AI reviewer's opinion of the finished work, and the count beside it is how many pieces that opinion covers. It only grades work that was actually produced, so a model that fails often is judged on its successes alone — 6.5/10 (18 pieces) is a weaker claim than 6.5/10 (26 pieces). Treat it as colour next to the pass rate, never as a substitute for it.
Performance shows two measured speeds on the machine named above: output speed is how fast the model writes its answer, and prefill speed is how fast it takes in your prompt and documents. Both matter for how a gezel feels — prefill governs the pause before it starts, while output governs how fast text appears. A dash means that round did not record a throughput probe.
Context is the working memory the model was given for these runs — how much it can hold at once. Memory used is the peak RAM the model and its engine actually occupied, which is the number to check against your own machine. When the engine reported its KV-cache precision, it appears beside the model name — for example, (kv: q8_0).
Each test round keeps its General capability and Office and knowledge work tables together under one provenance stamp. Earlier rounds stay separate from the latest because a change to gezel or to the task set can move a score without any model changing.
Size is the model's parameter count where we know it. Bigger is often but not always better: on office work in particular, some smaller models beat larger ones, and a model family's habits matter more than its size.
Choosing from this
A high score on the office set is the better guide for everyday document, planning, and analysis work. A high score on the general set matters more if you want a gezel writing or fixing code.
If a model you're considering isn't listed, it hasn't been measured here yet — which is not a verdict on it either way. The Models catalogue in the app will still tell you whether it fits this machine.