Raw, sourced benchmark data with full methodology — tok/s, memory, accuracy — plus downloadable CC-BY datasets you can cite.
The measured reverse-engineering: ~45% on SWE-bench Verified (18/40 blind), mid-50s to low-60s with local engineering, and a genuine 15–20-point wall. Where it matches, the 3 hard walls, and the biggest lever.
DatasetTok/s and memory data for local AI on Apple Silicon in 2026: Outlier measured, Ollama and LM Studio from public benchmarks.
DatasetA reference table mapping Mac unified RAM to which local AI models fit, their disk sizes, and measured tokens/sec on an M1 Ultra. Plan your build or buy.
DatasetRaw side-by-side data: 54 prompts sent to Outlier Core 27B and Claude Opus 4.7, scored by rubric. Outlier scored 98.9% overall, 100% on the 9 hardest.
DatasetMMLU for every shipping tier, measured through the shipped 1.11.804 build on an M1 Ultra: Nano 52.0%, Lite 77.0%, Quick 81.0%, Core 89.5%, Vision 3.8 82.5%. n=200, with the superseded 1.11.757 run kept for comparison.
DatasetRaw tok/s and memory data for Outlier's MoE engines on Plus 397B and Vision 35B. Why both streaming engines (V10, V11) were retired in favor of V9…
Free Nano + Lite. Pro is a one-time $249 and adds everything (Plus 397B included). Founders Lifetime is $249 once. Apple Silicon only.
Download for MacOpen weights ship constantly and most are not worth your disk. When one beats a tier Outlier already has, on hardware you already own, we say so and what it replaces. That is the only reason we email.
Your address and which of our sites you sent it from. No IP, no user agent, no referer. Nothing is sent until you press the button. Privacy.