Outlier  ›  how-to

How to compare AI models side by side on your Mac

Quick answer
  • Benchmarks tell you about average tasks. Compare mode tells you about your task.
  • Ask multiple tiers the same question in one go and read the answers next to each other.
  • It is the fastest way to find out whether you actually need the bigger, slower tier.
  • Runs locally, so comparing costs you time and nothing else.

The honest answer to "which model should I use?" is that it depends on what you're asking it. The dishonest answer is a single leaderboard number. Comparing on your own prompts settles it in about a minute.

Why a leaderboard can't answer this

A benchmark score is an average over someone else's task distribution. Your work is not that distribution. We publish MMLU per tier — Nano 53.0%, Lite 77.5%, Quick 81.5%, Core and Code 89.5%, all n=200 — and those numbers are real, but they cannot tell you whether Lite is good enough for the summarizing you do all day.

Two of our own tiers sit within the confidence interval of each other on MMLU while feeling quite different in use. The only way through that is to try both on something you actually care about.

Comparing

  1. Open compare mode from the header.
  2. Pick the tiers you're deciding between.
  3. Ask once. The prompt goes to each of them.
  4. Read them side by side and notice which one you'd actually ship.

Use a real task, not a riddle. The prompts that separate models in practice are the boring ones you repeat daily.

What you'll usually discover

Common questions

How do I know which local AI model is best for me?

Ask two or three the same real question and read the answers side by side. Benchmarks describe average tasks; compare mode describes yours.

Does comparing models cost extra?

No. Everything runs locally, so a comparison costs time and battery, not tokens.

Can I compare a big model and a small one?

Yes, and it is the most useful comparison there is — it tells you whether the extra RAM is buying you anything on the work you actually do.

Do I need all the models downloaded first?

Yes, each tier you want to compare has to be on disk, since inference runs on your machine.

Try it on your own Mac

One signed, notarized download. No account, no token bill, and it keeps working with the Wi-Fi off.

Get Outlier for Mac