Outlier  ›  data

Local AI and math word problems on a Mac: 12 of 20 right became 20 of 20

Quick answer
  • Outlier Lite, answering straight away, got 12 of 20 short word problems right. When it worked each one out first, out of sight, it got 20 of 20.
  • The price is time. The median reply took 8.0 seconds instead of 0.4, and the model wrote a median of 479 tokens of working instead of 3. No working showed up in the visible answer in any of the 20.
  • The app only does this for messages it reads as a word problem. Five ordinary questions took a median of 3.4 seconds. In 1.11.908 only Lite does it.

Measured on 2026-10-10 on a Mac Studio with an M1 Ultra and 64 GB, through the installed app. One machine, 20 problems, one run per setup (Table 2).

Outlier Pro: $249 once · or 4 × $62.25

Download free Buy Pro

Free to start (Nano and Lite). macOS 26+, Apple silicon. Refund window: 30 days, no questions asked.

I typed a shopping bill into our free Lite model: four pens at $1.25, three notebooks at $3.40, "give the amount only." It said $13.85. The bill is $15.20. Lite is the 9B model in the free download. A small model asked for just the number can get the number wrong. Below is how often it did, what I changed, and what the change costs.

How I tested it

I wrote 20 short word problems: a coffee order, a discount, a trip time, a bill split three ways, ages, a unit rate. Each ends with "Give the number only" or "Give the amount only", because that is how you ask when you want a figure to paste somewhere. Two of them: a train leaves at 9:40 and arrives at 13:15, how many minutes is the trip, and an $80 jacket is 15% off, what is the sale price.

I sent each one to the installed app's chat, the normal path with no special prompt, incognito, web search off, with up to 4,096 tokens allowed per reply. A reply counted as right if the last number in it was within a cent of the true answer. I also flagged any reply where the working showed up in the answer: a thinking tag, a phrase like "the user is asking", or a reply that opened with "Okay,". Then I asked 5 ordinary questions, such as "What is the capital of Canada?", to see whether anything else changed.

The "before" run is Outlier 1.11.905. The "after" run is the installed 1.11.908, today's download. Lite is a 9B model and the RAM guide lists it as fitting a 16 GB Mac.

What it got wrong

Straight away, Lite got 12 of 20 right. These are the 8 it missed in that run.

Table 1: the 8 misses on 1.11.905

Problem (shortened)Right answerLite said
4 pens at $1.25 and 3 notebooks at $3.40, total$15.20$13.85
5 workers take 12 days to build a wall, days for 6 workers106.25
Sam is 3 times as old as Lee; in 5 years Sam is twice as old; Lee's age now510
6.5 liters per 100 km, liters for 340 km22.16.5
3 coffees at $4.75, paid with a $20 bill, change$5.75$3.25
Walking 4.5 km per hour, minutes for 3 km4038
Tickets $12 adult, $7.50 child; 2 adults and 3 children$46.50$39
28 students, 3/7 are boys, how many girls1614

What I changed

My first try was the obvious one: switch on the model's own thinking by typing /think. On a different set of 10 terse problems, on 1.11.905, Lite went from 5 of 10 right to 9 of 10. But the median reply took 62.5 seconds and wrote 3,064 tokens, and 4 of the 10 replies had the working inside them. The app's standing instructions told the model to show no working at all, and the model was fighting them.

So I changed the instruction for one case. When the app reads a message as a word problem (numbers, an operation, a question), Lite is allowed to work it out privately, and the answer itself has to stay clean. Every other message is handled as before. That is what 1.11.908 does.

Table 2: Outlier Lite on the same 20 problems

SetupRightMedian time to finishMedian tokens writtenWorking in the answer
Straight answer (1.11.905)12 of 200.4 s30 of 20
Works it out first, out of sight (1.11.908)20 of 208.0 s4790 of 20

Why thinking first helps, and why my Nano page says the opposite

A model writes a small piece of text at a time and cannot go back. Ask for the number only and the first thing it writes is the answer. It has to produce "15.20" without having added anything up. A median reply of 3 tokens is exactly that. With room to work, it can write 4 x 1.25 = 5 and carry on. That is the usual explanation and these numbers fit it. I did not test another one.

I have a page that found the opposite on Nano: with thinking switched off, a grade-school maths set went from 33 of 60 to 46 of 60 right. The difference was the budget. There the model had 400 tokens, and its built-in preamble used 200 to 300 of them before it reached the sum. Here the working has up to 4,096 tokens and never reaches the screen. So thinking is not good or bad. It helps when it has room, and when it is kept out of the answer. See also what AI reasoning is.

What I did not ship

My first build did this for Nano and Lite together. On that build Nano got 15 of 20, took a median of 27.6 seconds, wrote a median of 810 tokens, and 2 of the 20 answers had working in them. I did not publish that build. In 1.11.908 Nano answers word problems the old way, and Core and Vision 3.8 are unchanged.

How far to trust 20 of 20

Not very far, and I would rather say so. It is 20 problems, I wrote them, and each setup ran once, so every figure above is one draw and not a rate. With 20 of 20, the 95% lower bound on the true rate is about 83%. Read it as "much better than 12 of 20", not "always right". If the number is a bill you are about to pay, use a calculator.

If you use another local app, you can run my test in ten minutes. Take 10 of your own word problems. Ask each once with "number only", and once with "work it out step by step, then give the final number". Count the right ones. I have not tried that on other apps, so I do not know which way yours will go. In Outlier, add /no_think at the end of a message to skip the thinking step. Put it at the end, not the start: a message that starts with a slash is read as an app command.

Frequently asked questions

Why does a small local AI get math word problems wrong?

Asked for just the number, the model has to write the answer as its first words, with no working. In my test Outlier Lite replied in a median of 3 tokens and got 12 of 20 right. The misses included repeating a number printed in the question and answering a different question than the one asked (Table 1).

Does making a local model think first make it more accurate?

On Lite, yes: 12 of 20 became 20 of 20 when it worked each problem out first, out of sight. It does not hold everywhere. On Nano, with a 400-token budget, turning thinking off scored better, 46 of 60 against 33 of 60, because the thinking used the budget up before the sum. It depends on whether the working has room.

How long does the thinking take?

A median of 8.0 seconds for the whole reply on a Mac Studio with an M1 Ultra, against 0.4 seconds without it. A smaller Mac will take longer; I only measured the one machine. The app does this only for messages it reads as a word problem. Five ordinary questions took a median of 3.4 seconds.

Does Outlier show the thinking?

No. In 20 of 20 answers a pattern check found no thinking tag and no working in the visible reply. The check looks for tags, phrases like "the user is asking", and replies that start with "Okay,". I did not read all 20 by eye.

Which Outlier tiers think first on word problems?

Only Lite, in 1.11.908. Nano, Core and Vision 3.8 answer word problems the way they did before.

Receipts: Every number here is a measurement I took myself on a Mac Studio with an M1 Ultra and 64 GB, running the installed signed Outlier app, on 2026-10-10. Same 20 problems and same scoring each time. The before figures (12 of 20, 0.4 s, 3 tokens, and the 5 of 10 to 9 of 10 typed /think test) are from 1.11.905. The after figures (20 of 20, 8.0 s, 479 tokens, and the 5 ordinary questions) are from the installed 1.11.908. The Nano figures are from a build I did not publish. Each setup ran once, at the app's default sampling, so none of it is a rate. The 33 of 60 and 46 of 60 are from my Nano page. No vendor supplied any of it, and no comparison with another app is claimed. The figures are published under CC BY 4.0; the models are on HuggingFace.

Try Outlier free

Free: Nano + Lite. Pro: $249 once · or 4 × $62.25. macOS 26+. In the US, Klarna or Afterpay at checkout: four payments, two weeks apart. Refund window: 30 days, no questions asked.

Download free (290 MB) Buy Pro

New models worth running