Local AI and math word problems on a Mac: 12 of 20 right became 20 of 20
- Outlier Lite, answering straight away, got 12 of 20 short word problems right. When it worked each one out first, out of sight, it got 20 of 20.
- The price is time. The median reply took 8.0 seconds instead of 0.4, and the model wrote a median of 479 tokens of working instead of 3. No working showed up in the visible answer in any of the 20.
- The app only does this for messages it reads as a word problem. Five ordinary questions took a median of 3.4 seconds. In 1.11.908 only Lite does it.
Measured on 2026-10-10 on a Mac Studio with an M1 Ultra and 64 GB, through the installed app. One machine, 20 problems, one run per setup (Table 2).
Outlier Pro: $249 once · or 4 × $62.25 · founders price $124.50 while seats last
Free to start (Nano and Lite). macOS 26+, Apple silicon. Refund window: 30 days, no questions asked.
I typed a shopping bill into our free Lite model: four pens at $1.25, three notebooks at $3.40, "give the amount only." It said $13.85. The bill is $15.20. Lite is the 9B model in the free download. A small model asked for just the number can get the number wrong. Below is how often it did, what I changed, and what the change costs.
How I tested it
I wrote 20 short word problems: a coffee order, a discount, a trip time, a bill split three ways, ages, a unit rate. Each ends with "Give the number only" or "Give the amount only", because that is how you ask when you want a figure to paste somewhere. Two of them: a train leaves at 9:40 and arrives at 13:15, how many minutes is the trip, and an $80 jacket is 15% off, what is the sale price.
I sent each one to the installed app's chat, the normal path with no special prompt, incognito, web search off, with up to 4,096 tokens allowed per reply. A reply counted as right if the last number in it was within a cent of the true answer. I also flagged any reply where the working showed up in the answer: a thinking tag, a phrase like "the user is asking", or a reply that opened with "Okay,". Then I asked 5 ordinary questions, such as "What is the capital of Canada?", to see whether anything else changed.
The "before" run is Outlier 1.11.905. The "after" run is the installed 1.11.908, today's download. Lite is a 9B model and the RAM guide lists it as fitting a 16 GB Mac.
What it got wrong
Straight away, Lite got 12 of 20 right. These are the 8 it missed in that run.
Table 1: the 8 misses on 1.11.905
| Problem (shortened) | Right answer | Lite said |
|---|---|---|
| 4 pens at $1.25 and 3 notebooks at $3.40, total | $15.20 | $13.85 |
| 5 workers take 12 days to build a wall, days for 6 workers | 10 | 6.25 |
| Sam is 3 times as old as Lee; in 5 years Sam is twice as old; Lee's age now | 5 | 10 |
| 6.5 liters per 100 km, liters for 340 km | 22.1 | 6.5 |
| 3 coffees at $4.75, paid with a $20 bill, change | $5.75 | $3.25 |
| Walking 4.5 km per hour, minutes for 3 km | 40 | 38 |
| Tickets $12 adult, $7.50 child; 2 adults and 3 children | $46.50 | $39 |
| 28 students, 3/7 are boys, how many girls | 16 | 14 |
- The car answer, 6.5, is the rate printed in the question. The sum was never done.
- The age answer, 10, is how old Lee will be in five years. The question asked about now.
- The ticket answer, $39, is what 2 adults and 2 children would pay. There were 3 children.
- I cannot read a cause off the other five numbers, and I did not dig further.
What I changed
My first try was the obvious one: switch on the model's own thinking by typing /think. On a different set of 10 terse problems, on 1.11.905, Lite went from 5 of 10 right to 9 of 10. But the median reply took 62.5 seconds and wrote 3,064 tokens, and 4 of the 10 replies had the working inside them. The app's standing instructions told the model to show no working at all, and the model was fighting them.
So I changed the instruction for one case. When the app reads a message as a word problem (numbers, an operation, a question), Lite is allowed to work it out privately, and the answer itself has to stay clean. Every other message is handled as before. That is what 1.11.908 does.
Table 2: Outlier Lite on the same 20 problems
| Setup | Right | Median time to finish | Median tokens written | Working in the answer |
|---|---|---|---|---|
| Straight answer (1.11.905) | 12 of 20 | 0.4 s | 3 | 0 of 20 |
| Works it out first, out of sight (1.11.908) | 20 of 20 | 8.0 s | 479 | 0 of 20 |
- The time is from sending the message to the end of the reply, not to the first word. For the first-word wait see time to first token.
- The 479 tokens include the working you never see. That is where the 8 seconds go.
- On 1.11.908 the 5 ordinary questions took a median of 3.4 seconds and 144 tokens, with no working in any answer. I did not time those five on the old build, so I have no before and after for them.
Why thinking first helps, and why my Nano page says the opposite
A model writes a small piece of text at a time and cannot go back. Ask for the number only and the first thing it writes is the answer. It has to produce "15.20" without having added anything up. A median reply of 3 tokens is exactly that. With room to work, it can write 4 x 1.25 = 5 and carry on. That is the usual explanation and these numbers fit it. I did not test another one.
I have a page that found the opposite on Nano: with thinking switched off, a grade-school maths set went from 33 of 60 to 46 of 60 right. The difference was the budget. There the model had 400 tokens, and its built-in preamble used 200 to 300 of them before it reached the sum. Here the working has up to 4,096 tokens and never reaches the screen. So thinking is not good or bad. It helps when it has room, and when it is kept out of the answer. See also what AI reasoning is.
What I did not ship
My first build did this for Nano and Lite together. On that build Nano got 15 of 20, took a median of 27.6 seconds, wrote a median of 810 tokens, and 2 of the 20 answers had working in them. I did not publish that build. In 1.11.908 Nano answers word problems the old way, and Core and Vision 3.8 are unchanged.
How far to trust 20 of 20
Not very far, and I would rather say so. It is 20 problems, I wrote them, and each setup ran once, so every figure above is one draw and not a rate. With 20 of 20, the 95% lower bound on the true rate is about 83%. Read it as "much better than 12 of 20", not "always right". If the number is a bill you are about to pay, use a calculator.
If you use another local app, you can run my test in ten minutes. Take 10 of your own word problems. Ask each once with "number only", and once with "work it out step by step, then give the final number". Count the right ones. I have not tried that on other apps, so I do not know which way yours will go. In Outlier, add /no_think at the end of a message to skip the thinking step. Put it at the end, not the start: a message that starts with a slash is read as an app command.
Frequently asked questions
Why does a small local AI get math word problems wrong?
Asked for just the number, the model has to write the answer as its first words, with no working. In my test Outlier Lite replied in a median of 3 tokens and got 12 of 20 right. The misses included repeating a number printed in the question and answering a different question than the one asked (Table 1).
Does making a local model think first make it more accurate?
On Lite, yes: 12 of 20 became 20 of 20 when it worked each problem out first, out of sight. It does not hold everywhere. On Nano, with a 400-token budget, turning thinking off scored better, 46 of 60 against 33 of 60, because the thinking used the budget up before the sum. It depends on whether the working has room.
How long does the thinking take?
A median of 8.0 seconds for the whole reply on a Mac Studio with an M1 Ultra, against 0.4 seconds without it. A smaller Mac will take longer; I only measured the one machine. The app does this only for messages it reads as a word problem. Five ordinary questions took a median of 3.4 seconds.
Does Outlier show the thinking?
No. In 20 of 20 answers a pattern check found no thinking tag and no working in the visible reply. The check looks for tags, phrases like "the user is asking", and replies that start with "Okay,". I did not read all 20 by eye.
Which Outlier tiers think first on word problems?
Only Lite, in 1.11.908. Nano, Core and Vision 3.8 answer word problems the way they did before.
Try Outlier free
Free: Nano + Lite. Pro: $249 once · or 4 × $62.25; $124.50 for the first 25 · or 4 × $31.13. macOS 26+. In the US, Klarna or Afterpay at checkout: four payments, two weeks apart. Refund window: 30 days, no questions asked.
Download free (290 MB) Buy ProOn your phone?