Outlier  ›  vs

Gemma 4 vs Qwen 3.6 on a Mac: which one to run

Quick answer
  • They fit the same Macs. Gemma 4 26B-A4B is a 15.61 GB file and Qwen 3.6 27B is 15.13 GB, both at 4-bit, and each peaks near 16 GB while it answers. Plan on a 24 GB Mac for either.
  • Gemma is faster on long answers, two to three times in my two runs (43.6 and 69.0 tokens a second against 20.7 and 21.6). On a short reply, under about 85 tokens, Qwen finished first in my July test, because Gemma spends about 3.7 seconds before it starts.
  • Qwen is the better model for knowledge and for code in a project: 89.5% against 81.0% on the same 200 questions, and 23 of 50 real repository bugs against 0 of 50. On short coding problems they were level, 20 and 21 of 23.
  • If you will run one, run Qwen 3.6 27B. Add Gemma 4 if you write long drafts and want them sooner.

Measured on a Mac Studio with an M1 Ultra and 64 GB, in Outlier builds from April to August 2026 (dates in the tables). Where I have two measurements of the same thing, I show both. Sources are in the receipts at the bottom.

Outlier Pro: $249 once · or 4 × $62.25

Download free Buy Pro

Free to start (Nano and Lite). macOS 26+, Apple silicon. Refund window: 30 days, no questions asked.

I ship both of these in Outlier. Gemma 4 26B-A4B is the Quick tier and Qwen 3.6 27B is the Core tier, so I have run the same tests on both, on one Mac Studio. I already wrote a page for each, how to run Gemma 4 on a Mac and how to run Qwen3.6 27B on a Mac. This one is for choosing between them.

One thing to know before the numbers. This is a mixture-of-experts model against a dense one. Gemma stores 26 billion parameters and uses about 4 billion for each word. Qwen uses all 27 billion for each word. A lot of what follows, the speed most of all, comes from that design, and you would see it in any pair built that way. It is not a ruling on Google against Alibaba. Gemma 4 also comes in a 31B, and the Qwen 3.6 family has a 35B-A3B. I have not measured either, and they would pair up differently.

The two side by side

Gemma 4 26B-A4B (Outlier Quick)Qwen 3.6 27B (Outlier Core)
Made byGoogleAlibaba's Qwen team
DesignMixture of experts: 26 billion stored, about 4 billion used per wordDense: all 27 billion used per word
4-bit MLX file15.61 GB15.13 GB (text only, image weights stripped)
Memory while it answersPeaks near 16 GBPeaks near 16 GB
Smallest Mac Outlier lists16 GB24 GB
Context window256K native, 32K default in the app256K native, 32K default in the app
Hugging Face licence tagApache 2.0Apache 2.0
In OutlierQuick, a Pro tierCore, a Pro tier and the default for agent work

Two notes on the memory rows. Outlier lists 16 GB as Quick's minimum, and the site's RAM table puts it in the 24 GB row. The file is 15.61 GB, so on a 16 GB Mac it loads with nothing to spare, and I would not plan around that. Second, a long chat costs memory on top of the file. The app's own readout for the 32K default estimates about 2.5 GB of context memory with Quick loaded and 2.0 GB with Core. Those come from a formula, not a measurement, and Quick's formula assumes 10 of its 30 layers keep a full-length cache, which is a cautious assumption I did not check.

How fast each one is

Table 2 is decode speed, how fast the words arrive once it has started. M1 Ultra, 4-bit MLX, one request at a time.

RunGemma 4 26B-A4BQwen 3.6 27BGap
A. Gemma: median of 20 requests through the shipped app, 2026-08-24 (1.11.804). Qwen: the 2026-06-09 batch43.620.72.1 times
B. Controlled test: one prompt at two token budgets, 2026-07-27 (build 1.11.669)69.021.63.2 times
C. Qwen only: a plain mlx-lm loop in August, and older harnesses in my lognot run16.6, and 5.5, 7.8, 11.6, 31unreconciled

So Gemma is 43.6 tokens a second in one run and 69.0 in another, and Qwen landed between 16.6 and 21.6 in the three runs I compare, with older figures I never reconciled. The builds and the methods differ, and I have not worked out which one accounts for the gap on Gemma, so I am not going to pick one for you. Both put Gemma at two to three times Qwen. That fits the design: on a Mac, each word is mostly limited by how fast the chip can read the model's weights from memory, and Gemma reads about 4 billion parameters' worth per word where Qwen reads all 27 billion.

The catch on short replies

Run B also measured the fixed wait before the first counted token: 3.72 seconds for Gemma and 1.03 seconds for Qwen. Gemma reasons before it answers, which may be part of that. I did not split it out. A reply takes the fixed wait plus its length divided by the speed, and Table 3 does that sum for both.

Reply lengthGemma 4 (3.72 s + length / 69.0)Qwen 3.6 (1.03 s + length / 21.6)Sooner
20 tokens, one short sentence4.0 s2.0 sQwen
50 tokens4.4 s3.3 sQwen
85 tokens5.0 s5.0 stie
200 tokens, a paragraph or two6.6 s10.3 sGemma
500 tokens11.0 s24.2 sGemma
1,000 tokens, a long draft18.2 s47.3 sGemma

The two cross at about 85 tokens. A one-line answer comes back sooner from Qwen. A long draft comes back much sooner from Gemma: 1,000 tokens in about 18 seconds against about 47.

Treat Table 3 as a sketch. I computed it from the two measured numbers for each model, I did not time each length. The fixed wait came from a short prompt sent straight to the engine on build 1.11.669, and I have not re-measured it on 1.11.912. A new chat in the app adds Outlier's own system prompt, and I did not time Gemma's first word there. Time to first token on a Mac has the Qwen side of that wait.

How good each one is

Same Mac, same app, the same questions where I could manage it. Where I could not, the sample column says so.

TestSampleGemma 4 (Quick)Qwen 3.6 (Core)
MMLU, general knowledge. Same 200 questions, same run, thinking off (1.11.804, 2026-08-24)20081.0% (162), 95% interval 75.0 to 85.889.5% (179), 95% interval 84.5 to 93.0
MMLU, a different sample, scored by reading a letter from a chat reply30079.3% (238)not run
Short coding problems asked in chat (August)2321 right, 250 s in all20 right, 394 s in all
HumanEval functions, run for real (July)2019 right20 right
Grade-school math, GSM8K (July)3026 right27 right
HumanEval, the full set (Gemma in April, Qwen in August, different harnesses)16421 right (12.8%)156 right (95.1%)
Real repository bugs, SWE-bench Verified, blind, fixed sample of 50500 right, all 50 with no patch (1.11.788, 2026-08-17)23 right, 7 with no patch (1.11.757, 2026-08-09)
Real repository bugs, an earlier and different set30 and 400 of 30 (July)18 of 40 (2026-06-25)

General knowledge

On 200 identical questions, Qwen got 179 right and Gemma got 162. The 95% intervals, 84.5 to 93.0 and 75.0 to 85.8, touch, and a plain two-proportion test on the two counts gives p of about 0.02. I read that as Qwen ahead by several points, not by exactly 8.5. Gemma's separate 300-question run scored 79.3%, which lands in the same place, and the earlier build, 1.11.757, gave 81.5% for Gemma and 89.5% for Qwen. The tier benchmarks page has the method and every tier.

Short code and math

On 23 short coding problems asked in chat, Gemma got 21 and Qwen got 20, and Gemma finished in 250 seconds against 394. Their misses did not overlap, so between them they solved all 23. That is a small set. One problem either way is noise. July's smaller checks agree: functions run for real, 19 of 20 and 20 of 20, and 30 grade-school math problems, 26 and 27.

One number disagrees with all of that. An April run of the full 164-problem HumanEval scored Gemma at 21, or 12.8%, with a harness that pulls the code out of a fenced block, while Qwen scored 156 of 164 in August through the app. I do not believe a model that gets 19 of 20 and 21 of 23 is really at 13% on the full set, and I have not worked out which test is wrong. I am showing it because it is already on my Gemma page and you should see both. HumanEval is also from 2021, so a 2026 model has very likely seen it.

Real repositories

This is the big gap. Given a real bug in a real project and left to find the fix, Qwen fixed 23 of 50 in a blind SWE-bench Verified run. Gemma fixed none. Its patches were not wrong. In all 50 tries its agent never wrote one, so nothing was graded. An earlier run on a different set of 30 tasks gave 0 of 30, with 29 of those stalling without taking any action. I wrote up the pattern in chat coding vs agentic coding.

Read that as a result about Gemma inside Outlier's agent loop. It says the pairing does not work for agent coding in my harness. It does not say Gemma cannot do it in someone else's, and I have not tried another. The short-code scores above say the model can write code. That is why Core is the default for agent work. These runs were on builds 1.11.757 and 1.11.788, and I have not re-run them on 1.11.912.

Which one I would run

Both are Pro tiers in Outlier. You switch between them from the model menu at the top of the window, and the free tiers are Nano and Lite. If you would rather run them without Outlier, the commands are on the Gemma 4 and Qwen3.6 27B pages. For Gemma against Llama, see Gemma vs Llama, and for Qwen against Llama, Qwen vs Llama.

What I did not test

Frequently asked questions

Is Gemma 4 or Qwen 3.6 better?

Qwen 3.6 27B scored higher on general knowledge (89.5% against 81.0% on the same 200 questions) and fixed 23 of 50 real repository bugs where Gemma 4 fixed none. On short coding problems they were level, 20 and 21 of 23. Gemma 4 26B is the faster one on long answers. Which is better depends on whether you wait on long drafts or on code in a project.

Which is faster on a Mac, Gemma 4 or Qwen 3.6?

Gemma 4 26B-A4B, for anything longer than about 85 tokens. In my two runs it wrote 43.6 and 69.0 tokens a second against 20.7 and 21.6 for Qwen 3.6 27B, on an M1 Ultra. For a one-line answer Qwen finished first in my July test, because Gemma takes about 3.7 seconds before its first counted token and Qwen about 1.0.

Which one needs more RAM?

About the same. The 4-bit Gemma 4 26B-A4B file is 15.61 GB and the Qwen 3.6 27B file is 15.13 GB, and each peaks near 16 GB while it answers. Outlier lists 16 GB as Gemma's minimum and 24 GB as Qwen's. I would put either on a 24 GB Mac.

Which is better for coding, Gemma 4 or Qwen 3.6?

For a single function, they are level in my tests. For fixing a bug in a real repository with an agent, Qwen 3.6 27B fixed 23 of 50 and Gemma 4 fixed 0 of 50 in Outlier's agent loop, which is a result about that pairing, not a ceiling for Gemma in every tool.

Can I use Gemma 4 and Qwen 3.6 commercially?

The Hugging Face repos I read on 2026-10-10 are tagged Apache 2.0: Qwen/Qwen3.6-27B, mlx-community/gemma-4-26b-a4b-it-4bit and Outlier-Ai/Outlier-Quick-26B-MLX-4bit. Outlier's own catalog entry for Quick still says Gemma Terms of Use, which does not match, and I have not reconciled the two. Read the model card of the copy you download before you build a product on it.

Receipts: Hardware: a Mac Studio with an M1 Ultra and 64 GB, 4-bit MLX, one request at a time. Speed run A: Qwen's 20.7 is the 2026-06-09 batch used across this site; Gemma's 43.6 is the median of 20 runs through the shipped 1.11.804 app on 2026-08-24. Speed run B (2026-07-27, build 1.11.669): the same prompt at two token budgets, 64 and 320, using the token counts the engine reported, with the slope giving decode speed and the intercept giving the fixed wait. A same-day split of that first pass gave Gemma 67.1. Qwen's 16.6 is a plain mlx-lm loop in August; the older Qwen figures of 5.5, 7.8, 11.6 and 31 come from other harnesses and I never reconciled them. Table 3 is arithmetic on run B, not a timing. File sizes, minimum memory and the 256K and 32K context figures are read from the tier catalog in Outlier 1.11.912; the 16 GB peaks are the generation footprint in the site's tier table; the 2.5 GB and 2.0 GB context estimates are the app's formula at 32K (8,192 bytes per token per full-length layer for 10 layers, and 4,096 for 16). MMLU: n=200, stratified across all 57 subjects, zero-shot, greedy, thinking off, through the shipped 1.11.804 app on 2026-08-24, scored by extracting the answer letter, the same 200 questions for both tiers; the 95% intervals are Wilson intervals computed from the counts. Gemma's 79.3% is n=300 across 50 of the 57 subjects, scored by reading a letter from a chat reply. Short coding: 23 chat prompts in August 2026. HumanEval n=20 and GSM8K n=30: greedy, July 2026, code run for real. Full HumanEval: Gemma in April, Qwen in the 1.11.804 run of 2026-08-24, different harnesses. Repository bugs: SWE-bench Verified, blind (no test patch, no file hint), 50 tasks from a fixed sample (seed 42), official Docker harness, no-patch tries counted as failures; Qwen on 1.11.757 and Gemma on 1.11.788, and the earlier sets were 40 and 30 tasks. Licence tags and the existence of the Gemma 4 31B and Qwen 3.6 35B-A3B models were read from the Hugging Face API on 2026-10-10. Every figure comes from one Mac Studio. None of it is a rate for your Mac.

Try Outlier free

Free: Nano + Lite. Pro: $249 once · or 4 × $62.25. macOS 26+. In the US, Klarna or Afterpay at checkout: four payments, two weeks apart. Refund window: 30 days, no questions asked.

Download free (290 MB) Buy Pro

New models worth running