Outlier  ›  how-to

How to run Qwen3.6 27B on a Mac: RAM, speed and where it falls short

Quick answer
  • Yes, Qwen3.6 27B runs on an Apple Silicon Mac with 24 GB or more. At 4-bit it is a 15.13 GB file and peaks near 16 GB of memory while it answers. A 16 GB Mac can't hold it.
  • On my M1 Ultra it writes about 20.7 tokens a second. Two later runs gave 21.6 and 16.6 (Table 3). That is roughly half the speed of Gemma 4 26B on the same Mac, because Qwen3.6 reads all 27 billion parameters for every word.
  • It scored the highest on general knowledge of any Outlier tier I measured, 89.5% on a 200-question sample. On real repository bugs it fixed 23 of 50.

Measured on a Mac Studio with an M1 Ultra and 64 GB. Speed is from my own runs, one request at a time. Quality is n=200 for knowledge questions, n=164 for HumanEval and n=50 for repository bugs. Sources and dates are in the receipts at the bottom.

Outlier Pro: $249 once · or 4 × $62.25

Download free Buy Pro

Free to start (Nano and Lite). macOS 26+, Apple silicon. Refund window: 30 days, no questions asked.

I ship Qwen3.6 27B inside Outlier as the Core tier, with the image weights stripped out so it is text only. I have timed it and tested it more than most people will. This page is what I measured on one Mac Studio, which Macs have room for it, three ways to run it, and where it falls short. If you only want the commands, jump to "Three ways to run it".

Which Mac has room for Qwen3.6 27B

Qwen3.6 27B is a dense model from Alibaba's Qwen team, released in April 2026. The Hugging Face model card is tagged Apache 2.0 (I read the tag on 2026-10-10). Dense means all 27 billion parameters are used for every word it writes. The original upload on Hugging Face is about 55.6 GB. At 4-bit in MLX format the text-only file Outlier ships is 15.13 GB. The community conversion on Hugging Face is 16.07 GB because it keeps the image-reading weights. While it answers, the model peaks near 16 GB of memory.

Table 1: Qwen3.6 27B at 4-bit, by Mac memory

Mac memoryLeft after the 16 GB peakMy read
16 GBnoneThe file alone is 15.13 GB. Outlier shows "Needs 24+ GB" on the Core row. Use Lite instead.
24 GBabout 8 GBThe listed minimum. Enough for macOS and a few apps at the default 32K context. I did not run a 24 GB Mac for this page.
32 GBabout 16 GBComfortable, with room for a browser, an editor and long chats.
64 GBabout 48 GBPlenty. This is the size of the Mac I measured on.

The middle column is just the memory minus 16 GB. For the general rule, see how much RAM local AI needs. If you have a 16 GB Mac, running AI on a 16 GB Mac covers what fits. Lite (Qwen3.5-9B) is a 5.04 GB file and free.

Three ways to run it

1. With MLX, free, from the Terminal

This is the library Outlier uses to load Core. It needs a Python environment.

pip install -U mlx-lm
mlx_lm.generate --model Outlier-Ai/Outlier-Core-27B-MLX-4bit \
  --prompt "Explain unified memory in two sentences." --max-tokens 200

The first run downloads about 15.2 GB from Hugging Face. After that it loads from disk. For a back-and-forth chat, use mlx_lm.chat with the same --model. That repo is my text-only 4-bit file of Qwen3.6-27B, the same one the Outlier app downloads, and it is tagged Apache 2.0. Outlier ships mlx-lm 0.31. If your copy is older and won't load the model, update it first.

If you would rather use the community conversion, the repo is mlx-community/Qwen3.6-27B-4bit. It is bigger, 16.07 GB, because it also holds the weights that read images. The Qwen3.5/3.6 loader in mlx-lm 0.31 skips those weights when it loads a model (I read the code), so you should get text only.

2. In Outlier, as the Core tier

Core is Qwen3.6 27B as a 4-bit MLX file. It is part of Pro. The free tiers are Nano and Lite.

  1. Download and open Outlier.
  2. Click the model button at the top of the window (its tooltip says "Switch model") and pick Core. If it isn't on your Mac yet, the row shows a download arrow and its size, 15.13 GB.
  3. If your Mac has less memory than Core's 24 GB minimum, the row says "Needs 24+ GB". If you click it anyway, the app says Core is sized for 24+ GB Macs and, on Pro, asks before it downloads.

Once it is downloaded it answers with the Wi-Fi off.

3. In Ollama or LM Studio

I have not tested Qwen3.6 27B in either, so I won't give you a command I haven't run. Both run GGUF files, which are a different format from the MLX file above. Search the model name in LM Studio's model browser, or look for it in Ollama's library. How to run Qwen on a Mac walks through both apps, and Ollama vs LM Studio helps you pick.

How fast it is

Table 2: decode speed on an M1 Ultra, 4-bit MLX, one request at a time

TierModelUsed per wordFileTokens a second
NanoQwen3.5-4B4B2.37 GB71.7
LiteQwen3.5-9B9B5.04 GB53.4
QuickGemma 4 26B-A4B4B15.61 GB43.6
CoreQwen3.6-27B27B15.13 GB20.7

The rows were measured on different days and builds, so read the ratios, not the third digit. The full reference table, with the method, is on Mac RAM to AI model size. A token is a word or part of one, and 20 a second is faster than most people read.

Why is Core about half as fast as Quick when the files are nearly the same size? On a Mac, writing each word is mostly limited by how quickly the chip can read the model's weights from memory. Quick is a mixture-of-experts model and reads about 4 billion parameters' worth per word. Core is dense and reads all 27 billion. My Gemma 4 page has the other side of that trade.

Table 3: three runs of Core on the same Mac

WhenHowTokens a second
2026-06-09Launch batch, one request at a time20.7
Late July 2026Bench script with a short and a long reply, so the fixed start-up cost drops out21.6
August 2026Plain mlx-lm loop, no app (Vision 3.8 did 15.5 in the same loop)16.6

These disagree by about a quarter and I have not reconciled them. My research log also holds earlier Core figures of 5.5, 7.8, 11.6 and 31 tokens a second from other harnesses, which I never reconciled either. The 20.7 is the figure on the rest of this site. If a friend asked, I would say "somewhere between 17 and 22 on an M1 Ultra", because three runs landed there.

Your Mac. Speed follows memory bandwidth. Apple lists 800 GB/s for the M1 Ultra and 273 GB/s for the M4 Pro, so 20.7 × 273 / 800 gives about 7 tokens a second on an M4 Pro. That is arithmetic, not a measurement, and a real chip can land on either side of it. The Core on an M4 Pro Mac mini page does the same sum.

The wait before the first word. Speed after the first word is one thing. On a new chat Core also reads Outlier's system prompt first. On 1.11.906 and again on 1.11.908, the first chat that had to build the saved part of that prompt waited 20.7 to 20.8 seconds for its first word on three of four test questions. Asking again started in 0.87 to 0.91 seconds on the same three. That is one machine and four questions. The full tables, and the part I could not explain, are on time to first token on a Mac.

How good it is

Table 4: Core (Qwen3.6 27B), my own tests

TestSampleCoreFor comparison
MMLU, general knowledge200 questions across all 57 subjects, thinking off89.5% (179 of 200), 95% interval about 4 points either wayVision 3.8 82.5%, Quick 81.0%, Lite 77.0%, same run
HumanEval, single functions16495.1% (156 of 164)Vision 3.8 155 of 164
Fix a real bug in a real repository (SWE-bench Verified, blind)50, fixed sample23 right, 7 with no patch at all, plus or minus 14 pointsQuick 0 of 30 on a separate set of 30

Knowledge. 89.5% is the highest score of any tier I measured. I did not measure Plus. At n=200 the interval is about 4 points near 90%, so read it as "high 80s". The same run put Vision 3.8, which is also a 27B model, at 82.5%. The tier benchmarks page has the method and every tier.

Short code. 156 of 164 sounds like a solved problem, and I don't think it is. HumanEval came out in 2021, so a 2026 model has very likely seen it. The number is useful for comparing tiers that faced the same set. Core and Vision 3.8 tied on a paired test (exact p = 1.000).

Real repositories. Given a bug in a real project and left to find the fix itself, Core fixed 23 of 50. An older build scored 18 of 40 on 2026-06-25, which is statistically the same. That is about half. It is a real result for a model that runs on your desk, and it is not Claude. I wrote up why in can a local 27B model code like Claude, and the chat-versus-agent split is in chat coding vs agentic coding.

When I would pick it

On a 24 GB or larger Mac, if you want one model for writing, explaining things and code, I would start with Core. If speed matters more than the last few points of quality, Quick answers about twice as fast and scored 81.0% to Core's 89.5%. If you need to read an image, that is Vision 3.8, since Core is text only. On a 16 GB Mac, start with Lite, which is free. If you aren't sure local AI is for you, download the free app and try Nano and Lite first.

What I did not test

Frequently asked questions

Can I run Qwen3.6 27B on a 16 GB Mac?

Not the 4-bit file I tested. It is 15.13 GB before macOS takes any memory, and it peaks near 16 GB while it answers. Outlier marks it Needs 24+ GB. Lite, a 5.04 GB 9B model, fits a 16 GB Mac and is free.

How much RAM does Qwen3.6 27B need on a Mac?

Outlier lists 24 GB as the minimum for the 4-bit file. The tier catalog records a peak of about 16 GB while it answers, and a plain mlx-lm run in August showed 15.45 GB. 32 GB is where there is comfortable headroom.

How fast is Qwen3.6 27B on Apple Silicon?

20.7 tokens a second on an M1 Ultra in my June 2026 batch. Two later runs with different harnesses gave 21.6 and 16.6. Slower chips will be slower, because a dense model's speed follows memory bandwidth. I only measured the M1 Ultra.

Is Qwen3.6 27B good for coding?

For single functions, yes: 156 of 164 on HumanEval, though that test is old enough that a 2026 model has probably seen it. For fixing a bug in a real repository it got 23 of 50 right in my blind SWE-bench Verified run.

Is Qwen3.6 27B free to run locally?

The weights are free to download, and the Hugging Face model card and the 4-bit repos are tagged Apache 2.0. The MLX route above costs nothing. In Outlier, Core is a Pro tier, while Nano and Lite are free.

Receipts: Speed: my own runs on a Mac Studio with an M1 Ultra and 64 GB, 4-bit MLX, one request at a time. The 20.7 is the 2026-06-09 batch, as on the RAM page. The 21.6 and 16.6 come from my research log: a late-July bench script that times a short and a long reply and subtracts the fixed start-up cost, and an August plain mlx-lm loop that also showed Core loading in 6.7 s and using 15.45 GB. The other tiers' speeds are from the 2026-06-09 batch and an August 2026 median of 20 for Quick. Core's file size and 24 GB minimum are read from the tier catalog in Outlier 1.11.911; the other tiers' sizes are from the site's tier table, read from the same catalog. The 16 GB peak is the generation footprint recorded in that table. The 55.6 GB original size and the 16.07 GB and 15.15 GB repo sizes were read from the Hugging Face API on 2026-10-10; the Apache 2.0 tags were read the same day. MMLU: n=200, stratified across all 57 subjects, greedy, thinking off, through the shipped 1.11.804 app on 2026-08-24, scored by extracting the answer letter. HumanEval: all 164 problems, same build and day, official tests run in a subprocess. Repository bugs: SWE-bench Verified, blind (no test patch, no file hint), 50 tasks from a fixed sample (seed 42), official Docker harness, the 7 that produced no patch counted as failures. First-word times: 1.11.906 and 1.11.908, four questions, 2026-10-10. The M4 Pro figure is 20.7 × 273 / 800, not a measurement. Every figure comes from the same Mac Studio and one chip. None of it is a rate for your Mac.

Try Outlier free

Free: Nano + Lite. Pro: $249 once · or 4 × $62.25. macOS 26+. In the US, Klarna or Afterpay at checkout: four payments, two weeks apart. Refund window: 30 days, no questions asked.

Download free (290 MB) Buy Pro

New models worth running