Local AI time to first token on a Mac: from 20 seconds to about 1
- On a new chat, Outlier Core and Vision 3.8 waited about 20 seconds for the first word on 1.11.905 (19.96 to 21.31 s across 16 requests). On 1.11.906 the same questions started in 0.87 to 0.88 seconds in 5 of 8 tests and in 1.58 to 2.04 seconds in the other 3.
- On 1.11.906 Nano and Lite start in about a third of a second on most questions: Nano 0.27 to 0.31 s, Lite 0.34 to 0.36 s, and Lite's coding question 0.72 s. On 1.11.905 they took 0.73 to 0.75 s and 1.16 to 1.20 s, and about 3.2 s and 5.8 s on the coding question.
- The catch: the first chat that has to build the saved part still waits 20.72 to 20.77 seconds on Core and Vision 3.8. Reuse did not change the words: on 1.11.906 the second ask matched the first word for word in 16 of 16 pairs.
Measured on 2026-10-10 on a Mac Studio with an M1 Ultra and 64 GB, through the installed app's engine. One machine, 4 questions per tier. I repeated it on 1.11.908, today's download, for Nano and Core, with the same result (Table 3).
Outlier Pro: $249 once · or 4 × $62.25 · founders price $124.50 while seats last
Free to start (Nano and Lite). macOS 26+, Apple silicon. Refund window: 30 days, no questions asked.
You press enter on a new chat and nothing happens for 20 seconds. That was the wait on Outlier Core and Vision 3.8 on 1.11.905. I timed it on my own Mac, changed one thing, and timed it again. Below are the request times from both runs, and the parts I could not explain.
Why a new chat waits at all
Before a model writes its first word it reads everything it was handed. That is your message plus the system prompt, the page of standing instructions an app sends with every chat. Reading is called prefill, and the wait it causes is the time to first token. It is a separate problem from how fast words appear afterwards.
A one-line question is about a dozen tokens. The system prompt is far longer. So most of the wait on a new chat is the app re-reading the same instructions it read in the last chat. A prefix cache saves that work so the next chat can skip it. The catch is that the saved work has to be exactly what a fresh read would have produced, and I learned that the hard way (see below).
There is a second reason a first message can be slow: the model is not in memory yet. This page does not measure that. The model was already loaded in every run.
The test
On the installed app, for each of four tiers, I asked four questions. For each question I turned the prefix cache off and on so nothing saved was left, asked it in a new one-message chat (ask 1), then asked the same question again in another new chat (ask 2). Temperature 0, so the same input should give the same words. The time is the engine's own first-token timer. I ran it on 1.11.905, then on 1.11.906.
Table 1: Outlier Core and Vision 3.8, first-token time
| Tier | Question | 905 ask 1 (s) | 905 ask 2 (s) | 906 ask 1 (s) | 906 ask 2 (s) |
|---|---|---|---|---|---|
| Core | Capital of Canada, one word | 20.00 | 19.98 | 20.72 | 0.87 |
| Core | Three commit-message tips | 21.28 | 21.31 | 2.03 | 2.04 |
| Core | Why the sky is blue, two sentences | 20.00 | 19.98 | 20.77 | 0.87 |
| Core | Python function: n-th Fibonacci number | 20.01 | 20.01 | 20.77 | 0.88 |
| Vision 3.8 | Capital of Canada, one word | 19.96 | 19.98 | 20.76 | 0.87 |
| Vision 3.8 | Three commit-message tips | 21.24 | 21.25 | 2.02 | 2.02 |
| Vision 3.8 | Why the sky is blue, two sentences | 19.96 | 19.97 | 20.73 | 0.87 |
| Vision 3.8 | Python function: n-th Fibonacci number | 21.24 | 21.24 | 1.59 | 1.58 |
- On 905 all 16 requests took 19.96 to 21.31 seconds. Ask 2 was no faster than ask 1, because nothing was saved between chats.
- On 906, ask 2 started in 0.87 to 0.88 seconds in 5 of 8 tests. In the other 3 it took 1.58 to 2.04 seconds. No ask 2 took longer than 2.04 s.
- On 906, ask 1 took 20.72 to 20.77 seconds in 5 of 8 tests. That is the cost of building the saved part, about the same as the whole wait on 905, paid once instead of every chat.
- In the other 3 tests ask 1 also took 1.59 to 2.03 seconds. I expected the cache toggle to leave nothing saved, so something was still there. I have not worked out why.
- Ask 2 is not zero because the date, any memory or project notes, and your message sit outside the saved part and are read every time.
Table 2: Outlier Nano and Lite, first-token time
| Tier | Question | 905 ask 1 (s) | 905 ask 2 (s) | 906 ask 1 (s) | 906 ask 2 (s) |
|---|---|---|---|---|---|
| Nano | Capital of Canada, one word | 0.73 | 0.75 | 2.53 | 0.27 |
| Nano | Three commit-message tips | 0.74 | 0.73 | 0.75 | 0.31 |
| Nano | Why the sky is blue, two sentences | 0.74 | 0.73 | 0.75 | 0.29 |
| Nano | Python function: n-th Fibonacci number | 3.24 | 3.23 | 3.75 | 0.29 |
| Lite | Capital of Canada, one word | 1.20 | 1.19 | 1.23 | 0.36 |
| Lite | Three commit-message tips | 1.16 | 1.18 | 1.20 | 0.34 |
| Lite | Why the sky is blue, two sentences | 1.16 | 1.18 | 1.19 | 0.36 |
| Lite | Python function: n-th Fibonacci number | 5.78 | 5.79 | 0.72 | 0.72 |
- Nano ask 2: 0.27 to 0.31 s on all four questions. On 905 it was 0.73 to 0.75 s on three of them and about 3.2 s on the coding question.
- Lite ask 2: 0.34 to 0.36 s on three questions and 0.72 s on the coding one. On 905 it was 1.16 to 1.20 s, and about 5.8 s on coding.
- Building the saved part can cost more than 905 did. Nano's ask 1 took 2.53 s on the first question (0.73 s on 905) and 3.75 s on the coding question (about 3.2 s on 905).
Table 3: the same test on 1.11.908, today's download (Core and Nano only)
| Tier | Question | 908 ask 1 (s) | 908 ask 2 (s) |
|---|---|---|---|
| Core | Capital of Canada, one word | 20.83 | 0.90 |
| Core | Three commit-message tips | 2.08 | 2.08 |
| Core | Why the sky is blue, two sentences | 20.83 | 0.91 |
| Core | Python function: n-th Fibonacci number | 20.83 | 0.91 |
| Nano | Capital of Canada, one word | 0.82 | 0.32 |
| Nano | Three commit-message tips | 0.77 | 0.31 |
| Nano | Why the sky is blue, two sentences | 0.79 | 0.31 |
| Nano | Python function: n-th Fibonacci number | 3.39 | 0.33 |
- Core ask 2: 0.90 to 0.91 s on three questions and 2.08 s on the commit-message one. Ask 1 took 20.83 s on those three and 2.08 s on the commit-message one. Same shape as 906.
- Nano ask 2: 0.31 to 0.33 s on all four.
- That run covered Nano and Core only, so Lite and Vision 3.8 have no 908 numbers here. It used the same test script as 906.
Every pair gave the same words on both asks, on every build: 16 of 16 on 905, 16 of 16 on 906 and 8 of 8 on 908.
Why I turned this off once
This is not the first time I saved the system prompt's work. In July I wrote on the system prompt page that it took the wait for a follow-up from about 18 seconds to about 1.3. At the start of October I ran the test I should have run first: does it change the answers?
I asked 10 questions on Lite, Core and Vision 3.8, through both the API route and the route the app window uses, twice, at temperature 0, with the saved system prompt on and then off. With it off, the two rounds matched each other 60 times out of 60. With it on, the answer matched the answer without it in 90 of 120. Thirty changed. One poem request opened with "The moon pulls tides across the sand" with the saved prompt and "The waves whisper secrets to the shore" without.
A faster first word is not worth a different answer, so I switched it off. That is why 1.11.905 waited 20 seconds. The July figure and these are different tests on different builds, so do not line them up.
What 1.11.906 does differently: the first chat builds the saved part with exactly the steps a later chat reuses, so there is one computation, not two to compare. I never proved why the old copy differed. My guess is it was built in a separate pass from a normal chat. I did not chase it down. I removed the difference instead and then checked the result, which is the 16 of 16 above.
Method and caveats
- Hardware: Mac Studio, M1 Ultra, 64 GB unified memory. One machine, one configuration.
- Software: the installed signed app, 1.11.905 first and then 1.11.906, in the early hours of 2026-10-10, Eastern time, and 1.11.908 later that morning. Not a dev build.
- Requests: sent to the app's local engine, one at a time, private chats, temperature 0, seed 7, up to 160 tokens, web search off. The app window sends a system prompt of its own, so what you see typing there can differ.
- The timer: the engine's first-token time, reported when the reply finishes. Printed to three decimals in the raw files and rounded to two here.
- Sample: 4 questions per tier, each asked twice, per build: 16 pairs on each of 905 and 906, 8 pairs on 908. Small, and the raw lines are above.
- Not covered: Outlier Quick and Plus, Lite and Vision 3.8 on 908, other Macs, an app restart, the speed of the rest of the reply, and the model-loading wait.
If you run a local model yourself
- Time the first word and the speed after it separately. They have different fixes. The speed-up guide covers both.
- Put the instructions that never change at the top and keep them identical from chat to chat. Anything that changes, like the date or memory, goes after them. A runtime can only save the part that stays put.
- Before you trust a saved prompt cache, ask the same question twice at temperature 0, with the cache and without, and compare the words. Mine differed 30 times in 120. The test took me about six minutes for each setting.
Questions about these first-token times
Why is the first message to a local AI slow?
Two different things can cause it. The model may not be in memory yet, which happens when a runtime unloads an idle model (Ollama's default is five minutes). Or the model has to read a long system prompt before the first word, which is prefill. This page measures the second one only. The model was already loaded in every run.
Does saving the system prompt's work change the answers?
It did once, and that is why I turned it off. In my October test the saved version gave different words in 30 of 120 pairs at temperature 0. In this test the second ask matched the first word for word in all 16 pairs on 1.11.906. In 3 of the 8 Core and Vision 3.8 pairs the first ask took about 2 seconds, which suggests it had also found a saved copy, so for those pairs the check shows two reused runs agree. The 5 pairs where the first ask took about 20.7 s are the ones that compare a reused copy against one built from scratch.
Will I get these times on my Mac?
This data cannot say. Every run was on one Mac Studio with an M1 Ultra and 64 GB. The seconds depend on your chip, because reading the system prompt is the part that got skipped. Core and Vision 3.8 are 27B models and both need 24 GB, so a smaller Mac would use Nano or Lite, which start in about a third of a second on most questions here.
Does this make the whole reply faster?
No. It only shortens the wait before the first word. I did not time the rest of the reply in this test, and nothing here suggests it changed.
Try Outlier free
Free: Nano + Lite. Pro: $249 once · or 4 × $62.25; $124.50 for the first 25 · or 4 × $31.13. macOS 26+. In the US, Klarna or Afterpay at checkout: four payments, two weeks apart. Refund window: 30 days, no questions asked.
Download free (290 MB) Buy ProOn your phone?