Outlier  ›  how-to

How to compare AI models side by side on your Mac

Quick answer
  • Benchmarks tell you about average tasks. Compare mode tells you about your task.
  • Ask multiple tiers the same question in one go and read the answers next to each other.
  • It is the fastest way to find out whether you actually need the bigger, slower tier.
  • Runs locally, so comparing costs you time and nothing else.

The honest answer to "which model should I use?" is that it depends on what you're asking it. The dishonest answer is a single leaderboard number. Comparing on your own prompts settles it in about a minute.

Why a leaderboard can't answer this

A benchmark score is an average over someone else's task distribution. Your work is not that distribution. We publish MMLU per tier — Nano 52.0%, Lite 77.0%, Quick 81.0%, Vision 3.8 82.5%, Core 89.5%, all n=200 — and those numbers are real, but they cannot tell you whether Lite is good enough for the summarizing you do all day.

Two of our own tiers sit within the confidence interval of each other on MMLU while feeling quite different in use. The only way through that is to try both on something you actually care about.

Where the leaderboard runs out, in our own numbers

Here is the lineup with the gap to the next tier down, because the gap is the part a single column hides. At n=200 the 95% interval is roughly ±7 points near 50% and ±4 near 90% — so a gap smaller than the interval is not a ranking, it is noise wearing a decimal point.

TierMMLU (n=200)Gap to the tier belowIs that gap bigger than the interval?
Core89.5%+7.0Yes — at the edge, near the ±4 end
Vision 3.882.5%+1.5No
Quick81.0%+4.0No — Lite and Quick are a statistical tie
Lite77.0%+25.0Yes, comfortably
Nano52.0%——

Read the fourth column rather than the second. Three of the tiers measured here are not separated from their neighbour by more than the measurement can see, and one gap — Nano to Lite — is so large that no comparison is needed to feel it. That is the whole case for comparing on your own prompts: for the tiers in the middle, the leaderboard genuinely does not know which is better for you, and it is not being coy about it.

Plus 397B is not in the table, and that is a fact about the measurement rather than the tier: the n=200 MMLU re-run covered five of the six shipping tiers and Plus was not one of them. A table that quietly dropped the sixth would be the kind of thing this page is arguing against.

Receipts: MMLU, n=200, deterministic stratified sample across 57 subjects, greedy (temperature 0, fixed seed), thinking off, on an M1 Ultra; all tiers re-run 2026-08-24 through the shipping v1.11.804 path so every figure comes from one regime on one build. Interval at that sample size is roughly ±7 points near 50% and ±4 near 90%. Gaps are arithmetic on those published figures, not a separate measurement.

Comparing

  1. Open compare mode from the header.
  2. Pick the tiers you're deciding between.
  3. Ask once. The prompt goes to each of them.
  4. Read them side by side and notice which one you'd actually ship.

Use a real task, not a riddle. The prompts that separate models in practice are the boring ones you repeat daily.

What you'll usually discover

Common questions

How do I know which local AI model is best for me?

Ask two or three the same real question and read the answers side by side. Benchmarks describe average tasks; compare mode describes yours.

Does comparing models cost extra?

No. Everything runs locally, so a comparison costs time and battery, not tokens.

Can I compare a big model and a small one?

Yes, and it is the most useful comparison there is — it tells you whether the extra RAM is buying you anything on the work you actually do.

Do I need all the models downloaded first?

Yes, each tier you want to compare has to be on disk, since inference runs on your machine.

Try it on your own Mac

One signed, notarized download. No account, no token bill, and it keeps working with the Wi-Fi off.

Free: Nano + Lite. Pro: $249 once · or 4 × $62.25. macOS 26+. In the US, Klarna or Afterpay at checkout: four payments, two weeks apart. Refund window: 30 days, no questions asked.

Download free (290 MB) Buy Pro

New models worth running