Outlier  ›  learn

Our free model and our paid model scored the same

Quick answer
  • Outlier Lite: 9B, 5.04 GB on disk, included free. MMLU 77.0% (154/200).
  • Outlier Quick: 26B, 15.61 GB on disk, part of Pro. MMLU 81.0% (162/200).
  • At n=200 the interval near those scores is roughly ±6 points. A 4-point gap sits inside it, so on this test the two are tied.
  • Three times the disk, three times the memory, no separation we can demonstrate on general knowledge.

We ran MMLU across every shipping tier in August, mostly to have honest numbers on the pricing page. The result we did not expect was that the free model and the first paid model landed on top of each other.

The numbers

Every measured tier, same run, same machine, same settings: Nano 52.0%, Lite 77.0%, Quick 81.0%, Core 89.5%, Vision 3.8 82.5%. When this was first written there was also a Code tier, and it scored 89.5% — identical to Core, because the two were the same weights with a different configuration. That is exactly why Code is no longer a separate tier: it now ships as Core, and the benchmark saying so was the argument for folding it in. The run was a control as much as a measurement, and it behaved.

The interesting pair is Lite and Quick. Lite is a 9B model in a 5 GB download. Quick is a 26B mixture-of-experts in a 15.6 GB download that needs 16 GB of memory rather than 12. On general knowledge, the extra 10 GB bought about four points, which at this sample size we cannot distinguish from noise.

What this does not mean

It does not mean Quick is pointless. MMLU is multiple-choice general knowledge. It says nothing about reasoning under a long context, about tool use, or about how a model behaves when it has to plan. We have measured cases where Quick clearly separates from smaller tiers, and one where it collapses to zero while a same-sized sibling works fine.

It does mean that if your work looks like MMLU, general questions with a known answer, the free tier may be all you need. We would rather you learned that from us than felt it later.

Why publish it

A benchmark that always ranks the expensive thing higher is marketing. This one did not, so here it is. It also gives us a real question to answer: whether Quick earns its slot on other evidence, or whether the lineup has one tier too many. We are measuring that rather than guessing.

Method

Every tier here runs on your own Mac, offline. Outlier ships Lite free. The FAQ explains what each tier is for, and support is one person.

New models worth running