Outlier  ›  data

Local AI time to first token on a Mac: from 20 seconds to about 1

Quick answer
  • On a new chat, Outlier Core and Vision 3.8 waited about 20 seconds for the first word on 1.11.905 (19.96 to 21.31 s across 16 requests). On 1.11.906 the same questions started in 0.87 to 0.88 seconds in 5 of 8 tests and in 1.58 to 2.04 seconds in the other 3.
  • On 1.11.906 Nano and Lite start in about a third of a second on most questions: Nano 0.27 to 0.31 s, Lite 0.34 to 0.36 s, and Lite's coding question 0.72 s. On 1.11.905 they took 0.73 to 0.75 s and 1.16 to 1.20 s, and about 3.2 s and 5.8 s on the coding question.
  • The catch: the first chat that has to build the saved part still waits 20.72 to 20.77 seconds on Core and Vision 3.8. Reuse did not change the words: on 1.11.906 the second ask matched the first word for word in 16 of 16 pairs.

Measured on 2026-10-10 on a Mac Studio with an M1 Ultra and 64 GB, through the installed app's engine. One machine, 4 questions per tier. I repeated it on 1.11.908, today's download, for Nano and Core, with the same result (Table 3).

Outlier Pro: $249 once · or 4 × $62.25

Download free Buy Pro

Free to start (Nano and Lite). macOS 26+, Apple silicon. Refund window: 30 days, no questions asked.

You press enter on a new chat and nothing happens for 20 seconds. That was the wait on Outlier Core and Vision 3.8 on 1.11.905. I timed it on my own Mac, changed one thing, and timed it again. Below are the request times from both runs, and the parts I could not explain.

Why a new chat waits at all

Before a model writes its first word it reads everything it was handed. That is your message plus the system prompt, the page of standing instructions an app sends with every chat. Reading is called prefill, and the wait it causes is the time to first token. It is a separate problem from how fast words appear afterwards.

A one-line question is about a dozen tokens. The system prompt is far longer. So most of the wait on a new chat is the app re-reading the same instructions it read in the last chat. A prefix cache saves that work so the next chat can skip it. The catch is that the saved work has to be exactly what a fresh read would have produced, and I learned that the hard way (see below).

There is a second reason a first message can be slow: the model is not in memory yet. This page does not measure that. The model was already loaded in every run.

The test

On the installed app, for each of four tiers, I asked four questions. For each question I turned the prefix cache off and on so nothing saved was left, asked it in a new one-message chat (ask 1), then asked the same question again in another new chat (ask 2). Temperature 0, so the same input should give the same words. The time is the engine's own first-token timer. I ran it on 1.11.905, then on 1.11.906.

Table 1: Outlier Core and Vision 3.8, first-token time

TierQuestion905 ask 1 (s)905 ask 2 (s)906 ask 1 (s)906 ask 2 (s)
CoreCapital of Canada, one word20.0019.9820.720.87
CoreThree commit-message tips21.2821.312.032.04
CoreWhy the sky is blue, two sentences20.0019.9820.770.87
CorePython function: n-th Fibonacci number20.0120.0120.770.88
Vision 3.8Capital of Canada, one word19.9619.9820.760.87
Vision 3.8Three commit-message tips21.2421.252.022.02
Vision 3.8Why the sky is blue, two sentences19.9619.9720.730.87
Vision 3.8Python function: n-th Fibonacci number21.2421.241.591.58

Table 2: Outlier Nano and Lite, first-token time

TierQuestion905 ask 1 (s)905 ask 2 (s)906 ask 1 (s)906 ask 2 (s)
NanoCapital of Canada, one word0.730.752.530.27
NanoThree commit-message tips0.740.730.750.31
NanoWhy the sky is blue, two sentences0.740.730.750.29
NanoPython function: n-th Fibonacci number3.243.233.750.29
LiteCapital of Canada, one word1.201.191.230.36
LiteThree commit-message tips1.161.181.200.34
LiteWhy the sky is blue, two sentences1.161.181.190.36
LitePython function: n-th Fibonacci number5.785.790.720.72

Table 3: the same test on 1.11.908, today's download (Core and Nano only)

TierQuestion908 ask 1 (s)908 ask 2 (s)
CoreCapital of Canada, one word20.830.90
CoreThree commit-message tips2.082.08
CoreWhy the sky is blue, two sentences20.830.91
CorePython function: n-th Fibonacci number20.830.91
NanoCapital of Canada, one word0.820.32
NanoThree commit-message tips0.770.31
NanoWhy the sky is blue, two sentences0.790.31
NanoPython function: n-th Fibonacci number3.390.33

Every pair gave the same words on both asks, on every build: 16 of 16 on 905, 16 of 16 on 906 and 8 of 8 on 908.

Why I turned this off once

This is not the first time I saved the system prompt's work. In July I wrote on the system prompt page that it took the wait for a follow-up from about 18 seconds to about 1.3. At the start of October I ran the test I should have run first: does it change the answers?

I asked 10 questions on Lite, Core and Vision 3.8, through both the API route and the route the app window uses, twice, at temperature 0, with the saved system prompt on and then off. With it off, the two rounds matched each other 60 times out of 60. With it on, the answer matched the answer without it in 90 of 120. Thirty changed. One poem request opened with "The moon pulls tides across the sand" with the saved prompt and "The waves whisper secrets to the shore" without.

A faster first word is not worth a different answer, so I switched it off. That is why 1.11.905 waited 20 seconds. The July figure and these are different tests on different builds, so do not line them up.

What 1.11.906 does differently: the first chat builds the saved part with exactly the steps a later chat reuses, so there is one computation, not two to compare. I never proved why the old copy differed. My guess is it was built in a separate pass from a normal chat. I did not chase it down. I removed the difference instead and then checked the result, which is the 16 of 16 above.

Method and caveats

If you run a local model yourself

Questions about these first-token times

Why is the first message to a local AI slow?

Two different things can cause it. The model may not be in memory yet, which happens when a runtime unloads an idle model (Ollama's default is five minutes). Or the model has to read a long system prompt before the first word, which is prefill. This page measures the second one only. The model was already loaded in every run.

Does saving the system prompt's work change the answers?

It did once, and that is why I turned it off. In my October test the saved version gave different words in 30 of 120 pairs at temperature 0. In this test the second ask matched the first word for word in all 16 pairs on 1.11.906. In 3 of the 8 Core and Vision 3.8 pairs the first ask took about 2 seconds, which suggests it had also found a saved copy, so for those pairs the check shows two reused runs agree. The 5 pairs where the first ask took about 20.7 s are the ones that compare a reused copy against one built from scratch.

Will I get these times on my Mac?

This data cannot say. Every run was on one Mac Studio with an M1 Ultra and 64 GB. The seconds depend on your chip, because reading the system prompt is the part that got skipped. Core and Vision 3.8 are 27B models and both need 24 GB, so a smaller Mac would use Nano or Lite, which start in about a third of a second on most questions here.

Does this make the whole reply faster?

No. It only shortens the wait before the first word. I did not time the rest of the reply in this test, and nothing here suggests it changed.

Receipts: Every number here is a measurement I took myself on a Mac Studio with an M1 Ultra and 64 GB, running the installed signed Outlier app, 1.11.905, 1.11.906 and 1.11.908, on 2026-10-10. The tables are the raw lines of my own test script, with the same four questions and settings on every build. The October comparison (90 of 120, 60 of 60) is from the same machine, at temperature 0, with the saved system prompt on and off. No vendor supplied any of it, and no comparison with another app is claimed. The figures are published under CC BY 4.0; the models are on HuggingFace.

Try Outlier free

Free: Nano + Lite. Pro: $249 once · or 4 × $62.25. macOS 26+. In the US, Klarna or Afterpay at checkout: four payments, two weeks apart. Refund window: 30 days, no questions asked.

Download free (290 MB) Buy Pro

New models worth running