How to run Qwen3.6 27B on a Mac: RAM, speed and where it falls short
- Yes, Qwen3.6 27B runs on an Apple Silicon Mac with 24 GB or more. At 4-bit it is a 15.13 GB file and peaks near 16 GB of memory while it answers. A 16 GB Mac can't hold it.
- On my M1 Ultra it writes about 20.7 tokens a second. Two later runs gave 21.6 and 16.6 (Table 3). That is roughly half the speed of Gemma 4 26B on the same Mac, because Qwen3.6 reads all 27 billion parameters for every word.
- It scored the highest on general knowledge of any Outlier tier I measured, 89.5% on a 200-question sample. On real repository bugs it fixed 23 of 50.
Measured on a Mac Studio with an M1 Ultra and 64 GB. Speed is from my own runs, one request at a time. Quality is n=200 for knowledge questions, n=164 for HumanEval and n=50 for repository bugs. Sources and dates are in the receipts at the bottom.
Outlier Pro: $249 once · or 4 × $62.25 · founders price $124.50 while seats last
Free to start (Nano and Lite). macOS 26+, Apple silicon. Refund window: 30 days, no questions asked.
I ship Qwen3.6 27B inside Outlier as the Core tier, with the image weights stripped out so it is text only. I have timed it and tested it more than most people will. This page is what I measured on one Mac Studio, which Macs have room for it, three ways to run it, and where it falls short. If you only want the commands, jump to "Three ways to run it".
Which Mac has room for Qwen3.6 27B
Qwen3.6 27B is a dense model from Alibaba's Qwen team, released in April 2026. The Hugging Face model card is tagged Apache 2.0 (I read the tag on 2026-10-10). Dense means all 27 billion parameters are used for every word it writes. The original upload on Hugging Face is about 55.6 GB. At 4-bit in MLX format the text-only file Outlier ships is 15.13 GB. The community conversion on Hugging Face is 16.07 GB because it keeps the image-reading weights. While it answers, the model peaks near 16 GB of memory.
Table 1: Qwen3.6 27B at 4-bit, by Mac memory
| Mac memory | Left after the 16 GB peak | My read |
|---|---|---|
| 16 GB | none | The file alone is 15.13 GB. Outlier shows "Needs 24+ GB" on the Core row. Use Lite instead. |
| 24 GB | about 8 GB | The listed minimum. Enough for macOS and a few apps at the default 32K context. I did not run a 24 GB Mac for this page. |
| 32 GB | about 16 GB | Comfortable, with room for a browser, an editor and long chats. |
| 64 GB | about 48 GB | Plenty. This is the size of the Mac I measured on. |
The middle column is just the memory minus 16 GB. For the general rule, see how much RAM local AI needs. If you have a 16 GB Mac, running AI on a 16 GB Mac covers what fits. Lite (Qwen3.5-9B) is a 5.04 GB file and free.
Three ways to run it
1. With MLX, free, from the Terminal
This is the library Outlier uses to load Core. It needs a Python environment.
pip install -U mlx-lm
mlx_lm.generate --model Outlier-Ai/Outlier-Core-27B-MLX-4bit \
--prompt "Explain unified memory in two sentences." --max-tokens 200
The first run downloads about 15.2 GB from Hugging Face. After that it loads from disk. For a back-and-forth chat, use mlx_lm.chat with the same --model. That repo is my text-only 4-bit file of Qwen3.6-27B, the same one the Outlier app downloads, and it is tagged Apache 2.0. Outlier ships mlx-lm 0.31. If your copy is older and won't load the model, update it first.
If you would rather use the community conversion, the repo is mlx-community/Qwen3.6-27B-4bit. It is bigger, 16.07 GB, because it also holds the weights that read images. The Qwen3.5/3.6 loader in mlx-lm 0.31 skips those weights when it loads a model (I read the code), so you should get text only.
2. In Outlier, as the Core tier
Core is Qwen3.6 27B as a 4-bit MLX file. It is part of Pro. The free tiers are Nano and Lite.
- Download and open Outlier.
- Click the model button at the top of the window (its tooltip says "Switch model") and pick Core. If it isn't on your Mac yet, the row shows a download arrow and its size, 15.13 GB.
- If your Mac has less memory than Core's 24 GB minimum, the row says "Needs 24+ GB". If you click it anyway, the app says Core is sized for 24+ GB Macs and, on Pro, asks before it downloads.
Once it is downloaded it answers with the Wi-Fi off.
3. In Ollama or LM Studio
I have not tested Qwen3.6 27B in either, so I won't give you a command I haven't run. Both run GGUF files, which are a different format from the MLX file above. Search the model name in LM Studio's model browser, or look for it in Ollama's library. How to run Qwen on a Mac walks through both apps, and Ollama vs LM Studio helps you pick.
How fast it is
Table 2: decode speed on an M1 Ultra, 4-bit MLX, one request at a time
| Tier | Model | Used per word | File | Tokens a second |
|---|---|---|---|---|
| Nano | Qwen3.5-4B | 4B | 2.37 GB | 71.7 |
| Lite | Qwen3.5-9B | 9B | 5.04 GB | 53.4 |
| Quick | Gemma 4 26B-A4B | 4B | 15.61 GB | 43.6 |
| Core | Qwen3.6-27B | 27B | 15.13 GB | 20.7 |
The rows were measured on different days and builds, so read the ratios, not the third digit. The full reference table, with the method, is on Mac RAM to AI model size. A token is a word or part of one, and 20 a second is faster than most people read.
Why is Core about half as fast as Quick when the files are nearly the same size? On a Mac, writing each word is mostly limited by how quickly the chip can read the model's weights from memory. Quick is a mixture-of-experts model and reads about 4 billion parameters' worth per word. Core is dense and reads all 27 billion. My Gemma 4 page has the other side of that trade.
Table 3: three runs of Core on the same Mac
| When | How | Tokens a second |
|---|---|---|
| 2026-06-09 | Launch batch, one request at a time | 20.7 |
| Late July 2026 | Bench script with a short and a long reply, so the fixed start-up cost drops out | 21.6 |
| August 2026 | Plain mlx-lm loop, no app (Vision 3.8 did 15.5 in the same loop) | 16.6 |
These disagree by about a quarter and I have not reconciled them. My research log also holds earlier Core figures of 5.5, 7.8, 11.6 and 31 tokens a second from other harnesses, which I never reconciled either. The 20.7 is the figure on the rest of this site. If a friend asked, I would say "somewhere between 17 and 22 on an M1 Ultra", because three runs landed there.
Your Mac. Speed follows memory bandwidth. Apple lists 800 GB/s for the M1 Ultra and 273 GB/s for the M4 Pro, so 20.7 × 273 / 800 gives about 7 tokens a second on an M4 Pro. That is arithmetic, not a measurement, and a real chip can land on either side of it. The Core on an M4 Pro Mac mini page does the same sum.
The wait before the first word. Speed after the first word is one thing. On a new chat Core also reads Outlier's system prompt first. On 1.11.906 and again on 1.11.908, the first chat that had to build the saved part of that prompt waited 20.7 to 20.8 seconds for its first word on three of four test questions. Asking again started in 0.87 to 0.91 seconds on the same three. That is one machine and four questions. The full tables, and the part I could not explain, are on time to first token on a Mac.
How good it is
Table 4: Core (Qwen3.6 27B), my own tests
| Test | Sample | Core | For comparison |
|---|---|---|---|
| MMLU, general knowledge | 200 questions across all 57 subjects, thinking off | 89.5% (179 of 200), 95% interval about 4 points either way | Vision 3.8 82.5%, Quick 81.0%, Lite 77.0%, same run |
| HumanEval, single functions | 164 | 95.1% (156 of 164) | Vision 3.8 155 of 164 |
| Fix a real bug in a real repository (SWE-bench Verified, blind) | 50, fixed sample | 23 right, 7 with no patch at all, plus or minus 14 points | Quick 0 of 30 on a separate set of 30 |
Knowledge. 89.5% is the highest score of any tier I measured. I did not measure Plus. At n=200 the interval is about 4 points near 90%, so read it as "high 80s". The same run put Vision 3.8, which is also a 27B model, at 82.5%. The tier benchmarks page has the method and every tier.
Short code. 156 of 164 sounds like a solved problem, and I don't think it is. HumanEval came out in 2021, so a 2026 model has very likely seen it. The number is useful for comparing tiers that faced the same set. Core and Vision 3.8 tied on a paired test (exact p = 1.000).
Real repositories. Given a bug in a real project and left to find the fix itself, Core fixed 23 of 50. An older build scored 18 of 40 on 2026-06-25, which is statistically the same. That is about half. It is a real result for a model that runs on your desk, and it is not Claude. I wrote up why in can a local 27B model code like Claude, and the chat-versus-agent split is in chat coding vs agentic coding.
When I would pick it
On a 24 GB or larger Mac, if you want one model for writing, explaining things and code, I would start with Core. If speed matters more than the last few points of quality, Quick answers about twice as fast and scored 81.0% to Core's 89.5%. If you need to read an image, that is Vision 3.8, since Core is text only. On a 16 GB Mac, start with Lite, which is free. If you aren't sure local AI is for you, download the free app and try Nano and Lite first.
What I did not test
- A 16 GB or 24 GB Mac, or any chip other than an M1 Ultra. The speed on an M4 will differ.
- Core with thinking turned on. The MMLU and HumanEval runs had it off.
- Chats longer than the 32K default context.
- Qwen3.6 27B in Ollama, LM Studio or llama.cpp.
- The community 16.07 GB conversion, in mlx-lm or for images.
- The two Terminal commands above. Outlier loads Core through the same mlx-lm library, but I did not retype them for this page.
- Any quantization other than 4-bit.
Frequently asked questions
Can I run Qwen3.6 27B on a 16 GB Mac?
Not the 4-bit file I tested. It is 15.13 GB before macOS takes any memory, and it peaks near 16 GB while it answers. Outlier marks it Needs 24+ GB. Lite, a 5.04 GB 9B model, fits a 16 GB Mac and is free.
How much RAM does Qwen3.6 27B need on a Mac?
Outlier lists 24 GB as the minimum for the 4-bit file. The tier catalog records a peak of about 16 GB while it answers, and a plain mlx-lm run in August showed 15.45 GB. 32 GB is where there is comfortable headroom.
How fast is Qwen3.6 27B on Apple Silicon?
20.7 tokens a second on an M1 Ultra in my June 2026 batch. Two later runs with different harnesses gave 21.6 and 16.6. Slower chips will be slower, because a dense model's speed follows memory bandwidth. I only measured the M1 Ultra.
Is Qwen3.6 27B good for coding?
For single functions, yes: 156 of 164 on HumanEval, though that test is old enough that a 2026 model has probably seen it. For fixing a bug in a real repository it got 23 of 50 right in my blind SWE-bench Verified run.
Is Qwen3.6 27B free to run locally?
The weights are free to download, and the Hugging Face model card and the 4-bit repos are tagged Apache 2.0. The MLX route above costs nothing. In Outlier, Core is a Pro tier, while Nano and Lite are free.
Try Outlier free
Free: Nano + Lite. Pro: $249 once · or 4 × $62.25; $124.50 for the first 25 · or 4 × $31.13. macOS 26+. In the US, Klarna or Afterpay at checkout: four payments, two weeks apart. Refund window: 30 days, no questions asked.
Download free (290 MB) Buy ProOn your phone?