Our free model and our paid model scored the same
- Outlier Lite: 9B, 5.04 GB on disk, included free. MMLU 77.0% (154/200).
- Outlier Quick: 26B, 15.61 GB on disk, part of Pro. MMLU 81.0% (162/200).
- At n=200 the interval near those scores is roughly ±6 points. A 4-point gap sits inside it, so on this test the two are tied.
- Three times the disk, three times the memory, no separation we can demonstrate on general knowledge.
We ran MMLU across every shipping tier in August, mostly to have honest numbers on the pricing page. The result we did not expect was that the free model and the first paid model landed on top of each other.
The numbers
Every measured tier, same run, same machine, same settings: Nano 52.0%, Lite 77.0%, Quick 81.0%, Core 89.5%, Vision 3.8 82.5%. When this was first written there was also a Code tier, and it scored 89.5% — identical to Core, because the two were the same weights with a different configuration. That is exactly why Code is no longer a separate tier: it now ships as Core, and the benchmark saying so was the argument for folding it in. The run was a control as much as a measurement, and it behaved.
The interesting pair is Lite and Quick. Lite is a 9B model in a 5 GB download. Quick is a 26B mixture-of-experts in a 15.6 GB download that needs 16 GB of memory rather than 12. On general knowledge, the extra 10 GB bought about four points, which at this sample size we cannot distinguish from noise.
What this does not mean
It does not mean Quick is pointless. MMLU is multiple-choice general knowledge. It says nothing about reasoning under a long context, about tool use, or about how a model behaves when it has to plan. We have measured cases where Quick clearly separates from smaller tiers, and one where it collapses to zero while a same-sized sibling works fine.
It does mean that if your work looks like MMLU, general questions with a known answer, the free tier may be all you need. We would rather you learned that from us than felt it later.
Why publish it
A benchmark that always ranks the expensive thing higher is marketing. This one did not, so here it is. It also gives us a real question to answer: whether Quick earns its slot on other evidence, or whether the lineup has one tier too many. We are measuring that rather than guessing.
Method
- MMLU, n=200, deterministic stratified sample across all 57 subjects, greedy decoding at temperature 0 with a fixed seed, thinking off, run through the shipping app path on build 1.11.757. Machine: M1 Ultra. Date: 2026-08-09.
- Confidence: roughly ±7 points near 50%, ±4 near 90%, at n=200.
- Cross-checked against an independent n=50 run on a different harness a fortnight earlier, which agreed within one point on every tier we could compare.
- The harness was validated against a known value before the campaign: Nano at n=25 returned 56% against a recorded 52%. A result near 25% would have meant broken routing rather than a weak model.
Every tier here runs on your own Mac, offline. Outlier ships Lite free. The FAQ explains what each tier is for, and support is one person.