Outlier tier benchmarks — MMLU on every shipping tier
Every shipping tier's MMLU score, measured through the app people actually download. The raw record is dataset.json (CC‑BY‑4.0).
MMLU, build 1.11.804, 2026-08-24
| Tier | MMLU | Correct |
|---|---|---|
| Nano | 52.0% | 104/200 |
| Lite | 77.0% | 154/200 |
| Quick | 81.0% | 162/200 |
| Core 27B | 89.5% | 179/200 |
| Vision 3.8 | 82.5% | 165/200 |
| Plus 397B | Not measured — see the dataset for why | |
How these were measured
Mounted the shipped 1.11.804 DMG, asserted /health reported 1.11.804, and queried that sidecar. Every row records the model that produced it and every run recorded zero transport errors.
Deterministic stratified sample across all 57 MMLU subjects (round-robin by subject, no RNG, reproducible). Zero-shot: question + A-D choices, answer with the letter only. Scored by letter extraction; the scorer self-validates against 6 known cases before every run. Requests go through the shipping app's /chat path so the measurement exercises the same code users run.
Decoding: greedy, temperature 0, fixed seed, thinking disabled. Sample: n = 200.
What changed, and why the numbers moved
Re-measured because three of these figures had no surviving results file. nano, lite and quick each came in slightly BELOW the 2026-08-09 run (1.0, 0.5 and 0.5 points). Sampling noise scatters both ways, so a one-sided shift suggests a small systematic difference between builds 1.11.757 and 1.11.804. Every difference is far inside the stated interval. The retired Code tier is dropped: it was Core's weights under a second name and never an independent measurement.
| Tier | MMLU on 1.11.757 | Correct |
|---|---|---|
| Nano | 53.0% | 106/200 |
| Lite | 77.5% | 155/200 |
| Quick | 81.5% | 163/200 |
| Core 27B | 89.5% | 179/200 |
HumanEval, build 1.11.804, 2026-08-24
| Tier | pass@1 | Passed |
|---|---|---|
| Core 27B | 95.1% | 156/164 |
| Vision 3.8 | 94.5% | 155/164 |
Full HumanEval, all 164 problems. The model is asked to complete the function; the completion is assembled with the problem's own test suite and check(entry_point) and executed in a subprocess with a 10s timeout. A solution counts only if the official tests pass. The scorer self-validates before any run: a known-correct completion must pass AND a known-wrong one must fail.
These two are tied. Core and Vision 3.8 were given the IDENTICAL 164 problems, so the right test is paired, not a comparison of two independent percentages. McNemar on the discordant pairs: Core solved 4 that 3.8 missed, 3.8 solved 3 that Core missed, 152 both, 5 neither. Exact two-sided p = 1.000. On single-function coding these two models are not distinguishable at this sample size.
An independent cross-check
An earlier independent run at n=50, using a different harness, gave nano 52, quick 82 and core 90 - within one point of these results on every tier.
New models worth running
Open weights ship constantly and most are not worth your disk. When one beats a tier Outlier already has, on hardware you already own, we say so and what it replaces. That is the only reason we email.
Your address and which site you sent it from. No IP, no user agent, no referer. Nothing is sent until you press the button. Privacy.