Outlier  ›  data

Why your local AI gives the same answer every time, and a 2-minute test

Quick answer
  • If a local model gives the same text every time at a temperature above 0 and no setting explains it, the model may be fine and the step that picks each word stuck on one random number. That was our bug, and a seed did not help because it never reached that step.
  • That was our bug. On Outlier 1.11.908, at the default temperature of 0.7, Nano, Lite and Core each gave the identical reply 3 times out of 3 on all three of my test prompts. Seed 7 and seed 8 gave the same text.
  • 1.11.909 fixed it. Every prompt now gets at least two different replies in three asks, a seed repeats and a new seed differs, and decode speed held (Tables 2 and 3). To test your own app, ask one open question 3 times (steps below).

Measured on 2026-10-10 on a Mac Studio with an M1 Ultra and 64 GB, through the installed app. One machine, 3 prompts asked 3 times each on each of 3 tiers, before and after (Tables 1 to 3).

Outlier Pro: $249 once · or 4 × $62.25

Download free Buy Pro

Free to start (Nano and Lite). macOS 26+, Apple silicon. Refund window: 30 days, no questions asked.

I asked each of our three models for a coffee shop name, three times in a row. Nano said "Brew & Bloom" three times. Lite said "The Daily Grind" three times. Core said "The Daily Grind" three times. The temperature was 0.7, the app's default, which is meant to give you something different each time. I tried a two-line poem and a fun fact about octopuses next. Same result on all nine combinations. Below is what was wrong, what I changed, and a test you can run on any app that lets you set a temperature. I have not tested any other app, so I can't tell you whether yours has this.

What I ran

Three open prompts, the kind with many good answers: Write a two-line poem about the sea, Suggest a name for a coffee shop. Just the name, and Give me one fun fact about octopuses. Each one went to each tier (Nano, Lite and Core) three times, as three separate requests with no earlier messages, incognito, web search off. I did not send a temperature, so the app used its own default, 0.7. I counted how many of the three replies were different from each other, comparing the whole text.

The requests went straight to the app's local engine, not through the chat window. I also asked the poem with seed 7, with seed 7 again, and with seed 8. And I timed decode speed with a 300-word story request, three runs per tier, using the speed the engine reports for each reply. The "before" run is Outlier 1.11.908. The "after" run is 1.11.909, the current download.

Before: 9 of 9 identical

Table 1: Outlier 1.11.908, temperature 0.7. Different replies out of 3 asks

TierPoemCoffee shop nameOctopus factSeed 7, seed 7, seed 8
Nano111all three the same text
Lite111all three the same text
Core111all three the same text

A 1 means all three replies matched word for word. Every cell is a 1. The coffee shop names were "Brew & Bloom" on Nano and "The Daily Grind" on Lite and Core, every time. Seed 8 should have given something different from seed 7. It gave the same text.

Why it happened

A model does not write a word directly. At each step it scores every possible next word, and the app then draws one at random, weighted by those scores. Temperature controls how flat the weights are. The draw needs a random key, and if the key is the same at every step, the draw stops being random. The same scores give the same word, so the same question gives the same reply.

That is what I found in ours. We run on MLX. The build in the app uses MLX 0.31.2 and its mlx_lm 0.31 library for sampling. MLX keeps its random-number state per thread. The library's sampling step is compiled so that it reads that state, and it reads the state of whichever thread loaded the sampling code first. Our chat runs on a different thread. So on that thread every draw used the same key, and the same noise was applied at every step of every reply. A seed went into a different thread's state, which is why seed 7 and seed 8 matched.

Nothing looks wrong in any single reply. Each one is fluent and on topic. You only see it by asking twice. It matters more than it sounds if you sample several answers on purpose: best-of-N, majority voting and self-consistency all become N copies of one answer. A separate R&D run of mine had already seen 60 pairs of problems repeat token for token.

One limit on the cause. In the test suite's Python, which has MLX 0.32.1, the library's own draw was random from any thread, so the stuck key did not reproduce there. The seed problem did. I only know this cause for the versions named above.

What I changed

Each reply now makes its own random key before it starts. If the request carries a seed, the key starts from that seed. If not, it starts from 64 random bits from the operating system. The key is split once per token, so no two tokens in a reply share a draw, and no thread's hidden state is read. The top-p, min-p and top-k filters are the library's own, in the library's order. Temperature 0 still takes the single most likely word and is unchanged.

I wrote 10 thread tests for it. Under the MLX the app ships, the old code failed 8 of them and the new code passes all 10.

Table 2: Outlier 1.11.909, temperature 0.7. Different replies out of 3 asks

TierPoemCoffee shop nameOctopus factSeed 7, seed 7, seed 8
Nano3237 and 7 matched, 8 differed
Lite3337 and 7 matched, 8 differed
Core3237 and 7 matched, 8 differed

Table 3: decode speed, median of 3 runs, tokens per second

Tier1.11.9081.11.909
Nano88.795.3
Lite67.967.7
Core22.722.7

The fix did not cost speed. Nano read faster on the second set of runs (88.4 to 89.5 before, 94.8 to 95.4 after). I did not look into why and I would not credit the fix for it.

Questions about an image go through a separate path in the app. I asked one image question three times on 1.11.908 and got three different replies, so that path was never stuck. I did not look into why.

Test your own app in two minutes

This works on any local app or server that lets you set a temperature.

  1. Set a temperature above 0. Use 0.7 if the app lets you pick. At temperature 0 the same reply every time is correct. The model takes its single most likely word. My first-word timing page tests at temperature 0 for that reason.
  2. Pick open questions. A two-line poem, a name for a coffee shop, a fun fact. Not "what is the capital of France", which has one right answer.
  3. Ask each one 3 times, from scratch. A new chat or a fresh request with no earlier messages, so the input is identical each time.
  4. Count the different replies. Three prompts, nine replies. Identical on all three prompts is the bug. My healthy run had two prompts with only two different replies out of three, and that is normal.
  5. If the app takes a seed, send the poem with seed 7 twice and seed 8 once. Working looks like 7 and 7 matching and 8 differing. Ours on 1.11.908 had all three matching.

If your app has an OpenAI-style local endpoint, this script does steps 1 to 5. It is a stripped-down version of what I ran, which used Outlier's own endpoint. I have not run this exact script against other apps. Fill in the address and model name. If your server wants a key, add the header.

import json, urllib.request

URL = "http://localhost:PORT/v1/chat/completions"   # your server's address
MODEL = "your-model-name"
PROMPTS = ["Write a two-line poem about the sea.",
           "Suggest a name for a coffee shop. Just the name.",
           "Give me one fun fact about octopuses."]

def ask(prompt, **extra):
    body = {"model": MODEL, "temperature": 0.7, "max_tokens": 120,
            "messages": [{"role": "user", "content": prompt}], **extra}
    req = urllib.request.Request(URL, json.dumps(body).encode(),
                                 {"Content-Type": "application/json"})
    with urllib.request.urlopen(req) as r:
        return json.load(r)["choices"][0]["message"]["content"].strip()

for p in PROMPTS:
    replies = [ask(p) for _ in range(3)]
    print(len(set(replies)), "different out of 3:", p)

a, b, c = (ask(PROMPTS[0], seed=s) for s in (7, 7, 8))
print("seed 7 twice matched:", a == b, "| seed 8 differed:", a != c)

Before you blame the sampler, rule out the dull causes. The temperature may really be 0 in the app, or a preset may override yours. A top-k of 1 is the same as temperature 0. A very small top-p does nearly the same. A cache in front of the model can hand back a stored reply. The test tells you the draw is stuck. It does not tell you why.

If you use Outlier, the current download has the fix (1.11.909 or newer). If you compare models, ask each question more than once. See how to compare models side by side, and how to run the local API if you want to point a script at it.

What I did not test

Frequently asked questions

Why does my local AI give the same answer every time?

Three common reasons. The temperature is 0, which is meant to repeat. The settings squeeze the choice down to one word, like a top-k of 1. Or the random draw is stuck, which was ours: on Outlier 1.11.908 the same question gave the same reply 3 times out of 3 on all nine prompt-and-tier pairs at temperature 0.7. The test above tells the first two from the third.

Why does temperature seem to do nothing on my local model?

In ours the temperature reached the sampler. The random key under it was stuck, so the same prompt always drew the same words. I did not run other temperatures, so I can't say what a different setting would have changed on 1.11.908.

Does a seed make a local model give the same answer twice?

On Outlier 1.11.909, seed 7 twice gave the same poem on all three tiers and seed 8 gave a different one. That is one prompt and one pair per tier on one Mac. On 1.11.908, seeds 7 and 8 gave the same text because the seed never reached the sampler.

Did this affect questions about images?

No. Image questions go through a separate path. I asked one image question three times on 1.11.908 and got three different replies.

Which Outlier version fixes it, and did the fix slow anything down?

1.11.909. Decode speed held on all three tiers: Nano 88.7 to 95.3 tokens per second, Lite 67.9 to 67.7, Core 22.7 to 22.7 (Table 3).

Receipts: Every number here is a measurement I took myself on a Mac Studio with an M1 Ultra and 64 GB, running the installed signed Outlier app, on 2026-10-10. Before figures are from 1.11.908, after figures from 1.11.909. Same three prompts, same method, default temperature 0.7, three asks per prompt per tier, incognito, web search off. Decode speed is the median of three 300-word story requests per tier. The cause comes from reading the library's code and from the thread tests I wrote for the fix. The image-question check is a single run of three asks. None of it is a rate, and I have not tested other apps.

Try Outlier free

Free: Nano + Lite. Pro: $249 once · or 4 × $62.25. macOS 26+. In the US, Klarna or Afterpay at checkout: four payments, two weeks apart. Refund window: 30 days, no questions asked.

Download free (290 MB) Buy Pro

New models worth running