Outlierdata

Outlier tier benchmarks — MMLU on every shipping tier

Every shipping tier's MMLU score, measured through the app people actually download. The raw record is dataset.json (CC‑BY‑4.0).

MMLU, build 1.11.804, 2026-08-24

TierMMLUCorrect
Nano52.0%104/200
Lite77.0%154/200
Quick81.0%162/200
Core 27B89.5%179/200
Vision 3.882.5%165/200
Plus 397BNot measured — see the dataset for why

How these were measured

Mounted the shipped 1.11.804 DMG, asserted /health reported 1.11.804, and queried that sidecar. Every row records the model that produced it and every run recorded zero transport errors.

Deterministic stratified sample across all 57 MMLU subjects (round-robin by subject, no RNG, reproducible). Zero-shot: question + A-D choices, answer with the letter only. Scored by letter extraction; the scorer self-validates against 6 known cases before every run. Requests go through the shipping app's /chat path so the measurement exercises the same code users run.

Decoding: greedy, temperature 0, fixed seed, thinking disabled. Sample: n = 200.

95% interval at n=200 is roughly +/-7 points near 50% and +/-4 points near 90%. Lite and Quick are a statistical tie.

What changed, and why the numbers moved

Re-measured because three of these figures had no surviving results file. nano, lite and quick each came in slightly BELOW the 2026-08-09 run (1.0, 0.5 and 0.5 points). Sampling noise scatters both ways, so a one-sided shift suggests a small systematic difference between builds 1.11.757 and 1.11.804. Every difference is far inside the stated interval. The retired Code tier is dropped: it was Core's weights under a second name and never an independent measurement.

TierMMLU on 1.11.757Correct
Nano53.0%106/200
Lite77.5%155/200
Quick81.5%163/200
Core 27B89.5%179/200

The original run. Kept because the figures above replaced it on the site. Superseded figures are kept rather than deleted, because a number that was published once should stay checkable.

HumanEval, build 1.11.804, 2026-08-24

Tierpass@1Passed
Core 27B95.1%156/164
Vision 3.894.5%155/164

Full HumanEval, all 164 problems. The model is asked to complete the function; the completion is assembled with the problem's own test suite and check(entry_point) and executed in a subprocess with a 10s timeout. A solution counts only if the official tests pass. The scorer self-validates before any run: a known-correct completion must pass AND a known-wrong one must fail.

These two are tied. Core and Vision 3.8 were given the IDENTICAL 164 problems, so the right test is paired, not a comparison of two independent percentages. McNemar on the discordant pairs: Core solved 4 that 3.8 missed, 3.8 solved 3 that Core missed, 152 both, 5 neither. Exact two-sided p = 1.000. On single-function coding these two models are not distinguishable at this sample size.

Read this figure carefully: HumanEval was published in 2021 and is very likely inside the training data of any 2026 model, so 95% partly measures memorisation rather than coding ability. We publish it because it is comparable across the tiers here - both faced the same contaminated set - not because it predicts how a model handles your codebase. The agentic SWE-bench figures are the harder and more honest test.

An independent cross-check

An earlier independent run at n=50, using a different harness, gave nano 52, quick 82 and core 90 - within one point of these results on every tier.

Receipts and limits: MMLU is multiple-choice general knowledge. It says nothing about reasoning over a long context, about tool use, or about how a model behaves when it has to plan. These are our own runs on our own hardware, not a third-party leaderboard, and they are published so they can be argued with.

New models worth running

Open weights ship constantly and most are not worth your disk. When one beats a tier Outlier already has, on hardware you already own, we say so and what it replaces. That is the only reason we email.

Your address and which site you sent it from. No IP, no user agent, no referer. Nothing is sent until you press the button. Privacy.