Gemma 4 vs Qwen 3.6 on a Mac: which one to run
- They fit the same Macs. Gemma 4 26B-A4B is a 15.61 GB file and Qwen 3.6 27B is 15.13 GB, both at 4-bit, and each peaks near 16 GB while it answers. Plan on a 24 GB Mac for either.
- Gemma is faster on long answers, two to three times in my two runs (43.6 and 69.0 tokens a second against 20.7 and 21.6). On a short reply, under about 85 tokens, Qwen finished first in my July test, because Gemma spends about 3.7 seconds before it starts.
- Qwen is the better model for knowledge and for code in a project: 89.5% against 81.0% on the same 200 questions, and 23 of 50 real repository bugs against 0 of 50. On short coding problems they were level, 20 and 21 of 23.
- If you will run one, run Qwen 3.6 27B. Add Gemma 4 if you write long drafts and want them sooner.
Measured on a Mac Studio with an M1 Ultra and 64 GB, in Outlier builds from April to August 2026 (dates in the tables). Where I have two measurements of the same thing, I show both. Sources are in the receipts at the bottom.
Outlier Pro: $249 once · or 4 × $62.25 · founders price $124.50 while seats last
Free to start (Nano and Lite). macOS 26+, Apple silicon. Refund window: 30 days, no questions asked.
I ship both of these in Outlier. Gemma 4 26B-A4B is the Quick tier and Qwen 3.6 27B is the Core tier, so I have run the same tests on both, on one Mac Studio. I already wrote a page for each, how to run Gemma 4 on a Mac and how to run Qwen3.6 27B on a Mac. This one is for choosing between them.
One thing to know before the numbers. This is a mixture-of-experts model against a dense one. Gemma stores 26 billion parameters and uses about 4 billion for each word. Qwen uses all 27 billion for each word. A lot of what follows, the speed most of all, comes from that design, and you would see it in any pair built that way. It is not a ruling on Google against Alibaba. Gemma 4 also comes in a 31B, and the Qwen 3.6 family has a 35B-A3B. I have not measured either, and they would pair up differently.
The two side by side
| Gemma 4 26B-A4B (Outlier Quick) | Qwen 3.6 27B (Outlier Core) | |
|---|---|---|
| Made by | Alibaba's Qwen team | |
| Design | Mixture of experts: 26 billion stored, about 4 billion used per word | Dense: all 27 billion used per word |
| 4-bit MLX file | 15.61 GB | 15.13 GB (text only, image weights stripped) |
| Memory while it answers | Peaks near 16 GB | Peaks near 16 GB |
| Smallest Mac Outlier lists | 16 GB | 24 GB |
| Context window | 256K native, 32K default in the app | 256K native, 32K default in the app |
| Hugging Face licence tag | Apache 2.0 | Apache 2.0 |
| In Outlier | Quick, a Pro tier | Core, a Pro tier and the default for agent work |
Two notes on the memory rows. Outlier lists 16 GB as Quick's minimum, and the site's RAM table puts it in the 24 GB row. The file is 15.61 GB, so on a 16 GB Mac it loads with nothing to spare, and I would not plan around that. Second, a long chat costs memory on top of the file. The app's own readout for the 32K default estimates about 2.5 GB of context memory with Quick loaded and 2.0 GB with Core. Those come from a formula, not a measurement, and Quick's formula assumes 10 of its 30 layers keep a full-length cache, which is a cautious assumption I did not check.
How fast each one is
Table 2 is decode speed, how fast the words arrive once it has started. M1 Ultra, 4-bit MLX, one request at a time.
| Run | Gemma 4 26B-A4B | Qwen 3.6 27B | Gap |
|---|---|---|---|
| A. Gemma: median of 20 requests through the shipped app, 2026-08-24 (1.11.804). Qwen: the 2026-06-09 batch | 43.6 | 20.7 | 2.1 times |
| B. Controlled test: one prompt at two token budgets, 2026-07-27 (build 1.11.669) | 69.0 | 21.6 | 3.2 times |
| C. Qwen only: a plain mlx-lm loop in August, and older harnesses in my log | not run | 16.6, and 5.5, 7.8, 11.6, 31 | unreconciled |
So Gemma is 43.6 tokens a second in one run and 69.0 in another, and Qwen landed between 16.6 and 21.6 in the three runs I compare, with older figures I never reconciled. The builds and the methods differ, and I have not worked out which one accounts for the gap on Gemma, so I am not going to pick one for you. Both put Gemma at two to three times Qwen. That fits the design: on a Mac, each word is mostly limited by how fast the chip can read the model's weights from memory, and Gemma reads about 4 billion parameters' worth per word where Qwen reads all 27 billion.
The catch on short replies
Run B also measured the fixed wait before the first counted token: 3.72 seconds for Gemma and 1.03 seconds for Qwen. Gemma reasons before it answers, which may be part of that. I did not split it out. A reply takes the fixed wait plus its length divided by the speed, and Table 3 does that sum for both.
| Reply length | Gemma 4 (3.72 s + length / 69.0) | Qwen 3.6 (1.03 s + length / 21.6) | Sooner |
|---|---|---|---|
| 20 tokens, one short sentence | 4.0 s | 2.0 s | Qwen |
| 50 tokens | 4.4 s | 3.3 s | Qwen |
| 85 tokens | 5.0 s | 5.0 s | tie |
| 200 tokens, a paragraph or two | 6.6 s | 10.3 s | Gemma |
| 500 tokens | 11.0 s | 24.2 s | Gemma |
| 1,000 tokens, a long draft | 18.2 s | 47.3 s | Gemma |
The two cross at about 85 tokens. A one-line answer comes back sooner from Qwen. A long draft comes back much sooner from Gemma: 1,000 tokens in about 18 seconds against about 47.
Treat Table 3 as a sketch. I computed it from the two measured numbers for each model, I did not time each length. The fixed wait came from a short prompt sent straight to the engine on build 1.11.669, and I have not re-measured it on 1.11.912. A new chat in the app adds Outlier's own system prompt, and I did not time Gemma's first word there. Time to first token on a Mac has the Qwen side of that wait.
How good each one is
Same Mac, same app, the same questions where I could manage it. Where I could not, the sample column says so.
| Test | Sample | Gemma 4 (Quick) | Qwen 3.6 (Core) |
|---|---|---|---|
| MMLU, general knowledge. Same 200 questions, same run, thinking off (1.11.804, 2026-08-24) | 200 | 81.0% (162), 95% interval 75.0 to 85.8 | 89.5% (179), 95% interval 84.5 to 93.0 |
| MMLU, a different sample, scored by reading a letter from a chat reply | 300 | 79.3% (238) | not run |
| Short coding problems asked in chat (August) | 23 | 21 right, 250 s in all | 20 right, 394 s in all |
| HumanEval functions, run for real (July) | 20 | 19 right | 20 right |
| Grade-school math, GSM8K (July) | 30 | 26 right | 27 right |
| HumanEval, the full set (Gemma in April, Qwen in August, different harnesses) | 164 | 21 right (12.8%) | 156 right (95.1%) |
| Real repository bugs, SWE-bench Verified, blind, fixed sample of 50 | 50 | 0 right, all 50 with no patch (1.11.788, 2026-08-17) | 23 right, 7 with no patch (1.11.757, 2026-08-09) |
| Real repository bugs, an earlier and different set | 30 and 40 | 0 of 30 (July) | 18 of 40 (2026-06-25) |
General knowledge
On 200 identical questions, Qwen got 179 right and Gemma got 162. The 95% intervals, 84.5 to 93.0 and 75.0 to 85.8, touch, and a plain two-proportion test on the two counts gives p of about 0.02. I read that as Qwen ahead by several points, not by exactly 8.5. Gemma's separate 300-question run scored 79.3%, which lands in the same place, and the earlier build, 1.11.757, gave 81.5% for Gemma and 89.5% for Qwen. The tier benchmarks page has the method and every tier.
Short code and math
On 23 short coding problems asked in chat, Gemma got 21 and Qwen got 20, and Gemma finished in 250 seconds against 394. Their misses did not overlap, so between them they solved all 23. That is a small set. One problem either way is noise. July's smaller checks agree: functions run for real, 19 of 20 and 20 of 20, and 30 grade-school math problems, 26 and 27.
One number disagrees with all of that. An April run of the full 164-problem HumanEval scored Gemma at 21, or 12.8%, with a harness that pulls the code out of a fenced block, while Qwen scored 156 of 164 in August through the app. I do not believe a model that gets 19 of 20 and 21 of 23 is really at 13% on the full set, and I have not worked out which test is wrong. I am showing it because it is already on my Gemma page and you should see both. HumanEval is also from 2021, so a 2026 model has very likely seen it.
Real repositories
This is the big gap. Given a real bug in a real project and left to find the fix, Qwen fixed 23 of 50 in a blind SWE-bench Verified run. Gemma fixed none. Its patches were not wrong. In all 50 tries its agent never wrote one, so nothing was graded. An earlier run on a different set of 30 tasks gave 0 of 30, with 29 of those stalling without taking any action. I wrote up the pattern in chat coding vs agentic coding.
Read that as a result about Gemma inside Outlier's agent loop. It says the pairing does not work for agent coding in my harness. It does not say Gemma cannot do it in someone else's, and I have not tried another. The short-code scores above say the model can write code. That is why Core is the default for agent work. These runs were on builds 1.11.757 and 1.11.788, and I have not re-run them on 1.11.912.
Which one I would run
- One model for everything, on a 24 GB or larger Mac: Qwen 3.6 27B.
- Code in a project folder, with the agent: Qwen 3.6 27B. Gemma got nothing in my harness.
- Long drafts, summaries and explanations, where the wait is what bothers you: Gemma 4. At 500 tokens my July sum has it at about 11 seconds against 24.
- Short questions and one-line answers: Qwen, for the shorter wait.
- A 16 GB Mac: neither, with room to spare. Start with Lite, which is free.
Both are Pro tiers in Outlier. You switch between them from the model menu at the top of the window, and the free tiers are Nano and Lite. If you would rather run them without Outlier, the commands are on the Gemma 4 and Qwen3.6 27B pages. For Gemma against Llama, see Gemma vs Llama, and for Qwen against Llama, Qwen vs Llama.
What I did not test
- Gemma 4 31B and Qwen 3.6 35B-A3B, which would pair up differently.
- Any chip other than an M1 Ultra, and any Mac with 16 or 24 GB.
- Either model in Ollama, LM Studio or llama.cpp.
- A paired test of thinking on against thinking off for these two. The MMLU runs had it off.
- Images, audio, and chats longer than the 32K default.
- Gemma's first-word wait on a new chat in the app.
- Gemma as an agent in any harness other than Outlier's.
Frequently asked questions
Is Gemma 4 or Qwen 3.6 better?
Qwen 3.6 27B scored higher on general knowledge (89.5% against 81.0% on the same 200 questions) and fixed 23 of 50 real repository bugs where Gemma 4 fixed none. On short coding problems they were level, 20 and 21 of 23. Gemma 4 26B is the faster one on long answers. Which is better depends on whether you wait on long drafts or on code in a project.
Which is faster on a Mac, Gemma 4 or Qwen 3.6?
Gemma 4 26B-A4B, for anything longer than about 85 tokens. In my two runs it wrote 43.6 and 69.0 tokens a second against 20.7 and 21.6 for Qwen 3.6 27B, on an M1 Ultra. For a one-line answer Qwen finished first in my July test, because Gemma takes about 3.7 seconds before its first counted token and Qwen about 1.0.
Which one needs more RAM?
About the same. The 4-bit Gemma 4 26B-A4B file is 15.61 GB and the Qwen 3.6 27B file is 15.13 GB, and each peaks near 16 GB while it answers. Outlier lists 16 GB as Gemma's minimum and 24 GB as Qwen's. I would put either on a 24 GB Mac.
Which is better for coding, Gemma 4 or Qwen 3.6?
For a single function, they are level in my tests. For fixing a bug in a real repository with an agent, Qwen 3.6 27B fixed 23 of 50 and Gemma 4 fixed 0 of 50 in Outlier's agent loop, which is a result about that pairing, not a ceiling for Gemma in every tool.
Can I use Gemma 4 and Qwen 3.6 commercially?
The Hugging Face repos I read on 2026-10-10 are tagged Apache 2.0: Qwen/Qwen3.6-27B, mlx-community/gemma-4-26b-a4b-it-4bit and Outlier-Ai/Outlier-Quick-26B-MLX-4bit. Outlier's own catalog entry for Quick still says Gemma Terms of Use, which does not match, and I have not reconciled the two. Read the model card of the copy you download before you build a product on it.
Try Outlier free
Free: Nano + Lite. Pro: $249 once · or 4 × $62.25; $124.50 for the first 25 · or 4 × $31.13. macOS 26+. In the US, Klarna or Afterpay at checkout: four payments, two weeks apart. Refund window: 30 days, no questions asked.
Download free (290 MB) Buy ProOn your phone?