Outlier  ›  how-to

How to run Gemma 4 on a Mac: RAM, speed and where it falls short

Quick answer
  • Yes, Gemma 4 runs on an Apple Silicon Mac. The 26B-A4B build (26 billion parameters, about 4 billion used per word) is a 15.61 GB file at 4-bit and peaks near 16 GB of memory while it answers. A 24 GB Mac is the comfortable size. A 16 GB Mac has nothing to spare.
  • On my M1 Ultra it writes 43.6 tokens a second, a little over twice the 20.7 of the 27B Qwen model I ship as Core (measured on different days, see Table 2).
  • It is a good chat model and a poor agent. It got 21 of 23 short coding problems right and 0 of 30 real repository bugs.

Measured on a Mac Studio with an M1 Ultra and 64 GB. Speed is the median of 20 runs through the shipped app bundle. Quality is n=300 for knowledge questions, n=23 for short code and n=30 for repository bugs. Sources and dates are in the receipts at the bottom.

Outlier Pro: $249 once · or 4 × $62.25

Download free Buy Pro

Free to start (Nano and Lite). macOS 26+, Apple silicon. Refund window: 30 days, no questions asked.

I ship Gemma 4 inside Outlier as the Quick tier, so I have timed it and tested it more than most people. This page is what I measured on one Mac Studio, which Macs have room for it, three ways to run it, and the two places it let me down. If you only want the commands, jump to "Three ways to run it".

Which Mac has room for Gemma 4

Gemma 4 comes in five sizes: E2B, E4B, 12B, 26B-A4B and 31B. The Gemma vs Llama page has the launch details and the licence. I measured one of them, the 26B-A4B. It is a mixture-of-experts model. The file holds 26 billion parameters, but only about 4 billion of them are used for each word it writes. At 4-bit in MLX format the file is 15.61 GB. While it answers, it peaks near 16 GB of memory.

Table 1: Gemma 4 26B-A4B at 4-bit, by Mac memory

Mac memoryLeft after the 16 GB peakMy read
16 GBabout 0 GBThis is the listed minimum. There is no spare memory for macOS and your apps, so I would not plan around it.
24 GBabout 8 GBFits with a browser and an editor open.
32 GBabout 16 GBFits, with room for long chats.
64 GBabout 48 GBPlenty of room.

The middle column is just the memory minus 16 GB. I did not time a 16 GB Mac for this page. For the general rule, see how much RAM local AI needs.

If you have 8 GB or 16 GB, the free tiers fit better. Outlier's Nano is a 2.37 GB file and Lite is 5.04 GB. They are Qwen models, not Gemma. If you want Gemma itself on a small Mac, E2B and E4B are the sizes Google makes for that. I have not measured them, so I can't give you a speed.

Three ways to run it

1. With MLX, free, from the Terminal

This is the same library Outlier uses to load Quick. It needs a Python environment.

pip install -U mlx-lm
mlx_lm.generate --model mlx-community/gemma-4-26b-a4b-it-4bit \
  --prompt "Explain unified memory in two sentences."

The first run downloads about 15.6 GB from Hugging Face. After that it loads from disk. For a back-and-forth chat, use mlx_lm.chat with the same --model. Outlier ships mlx-lm 0.31. If your copy is older and won't load the model, update it first. The Hugging Face repo is tagged Apache 2.0.

2. In Outlier, as the Quick tier

Quick is Gemma 4 26B-A4B as a 4-bit MLX file. It is part of Pro. The free tiers are Nano and Lite.

  1. Download and open Outlier.
  2. Open the model menu at the top of the chat and pick Quick. If it isn't on your Mac yet, you download it from that same menu (15.61 GB).
  3. If your Mac has less memory than Quick's 16 GB minimum, the menu warns you before it downloads anything.

Once it is downloaded it answers with the Wi-Fi off.

3. In Ollama or LM Studio

Google's own docs list both as ways to run Gemma 4. I have not tested Gemma 4 in either, so I won't give you a command that I haven't run. The steps are the same as for any other model: search for Gemma 4 in LM Studio's model browser, or pull it by name in Ollama. How to run Llama on a Mac walks through both apps, and Ollama vs LM Studio helps you pick.

How fast it is

Table 2: decode speed on an M1 Ultra, 4-bit MLX, one request at a time

TierModelUsed per wordFileTokens a secondMeasured
NanoQwen3.5-4B4B2.37 GB71.7June 2026 batch
LiteQwen3.5-9B9B5.04 GB53.4June 2026 batch
QuickGemma 4 26B-A4B4B15.61 GB43.62026-08-24, Outlier 1.11.804, median of 20
CoreQwen3.6-27B27B15.13 GB20.7June 2026 batch

The rows were measured on different days and builds, so read the ratios, not the third digit. The full reference table, with the method, is on Mac RAM to AI model size.

Why is Quick about twice as fast as Core when the files are the same size? On a Mac, writing each word is mostly limited by how quickly the chip can read the model's weights from memory. Quick reads about 4 billion parameters' worth per word. Core reads all 27 billion. Quick is still a little slower than the 9B Lite, which I put down to the routing work a mixture-of-experts model does on every word. I haven't isolated that.

One more thing about speed. Quick reasons before it answers. In my knowledge test it wrote about 146 reasoning tokens before each answer letter. At 43.6 tokens a second that is roughly 3 seconds of thinking before the first line shows up. That is arithmetic from two measured numbers, not a timing of its own. I did not time first-word latency on Quick.

How good it is

Table 3: Quick (Gemma 4 26B-A4B), my own tests

TestSampleQuickFor comparison
MMLU, general knowledge300 questions across 50 of 57 subjects79.3% (238 of 300), 95% interval 74.4% to 83.5%Lite 78.5%, same sprint
Short coding problems2321 right, 250 seconds in allCore 20 right, 394 seconds
Fix a real bug in a real repository (SWE-bench Verified, blind)300 rightCore 23 of 50 on the same harness
HumanEval16421 right (12.8%)not run for this page

Knowledge. On the MMLU sample, Quick and the 9B Lite are tied inside the noise. If all you need is facts and explanations, the free 5 GB Lite is a fair alternative. Quick's score came from reading an answer letter out of a chat reply, which is noisier than the standard scoring and can miss right answers. Lite's number is from the same sprint's table, and the two may not be scored the same way.

Short code. Asked for a single function with the problem already stated, Quick matched the 27B Core and finished in 37% less time. Between them they solved all 23, because their misses did not overlap. The sample is small, so one point either way is noise.

Real repositories. This is where it falls down. Given a bug in a real project and left to find the fix itself, Quick produced no usable patch in 29 of 30 tries. It stalled while searching and never settled on an edit. I wrote that test up in chat coding vs agentic coding. Core stays the default for agent work for that reason.

HumanEval. An older run in April scored Quick 21 of 164, with a harness that pulls the code out of a fenced block. That disagrees with the 21 of 23 above, and I haven't worked out why. Treat Quick's coding as short-function help, and use Core for anything larger.

When I would pick it

On a 24 GB or larger Mac, for quick answers, drafting and explaining things, Quick is a good fit, and it is the faster of the two big models. For code inside a project, use Core. On a 16 GB Mac, start with Lite, which is free. If you aren't sure local AI is for you, download the free app and try Nano and Lite first. They cost nothing.

What I did not test

Frequently asked questions

Can I run Gemma 4 on a MacBook Air?

The 26B-A4B is a 15.61 GB file that peaks near 16 GB, so a 16 GB Air has no room to spare and I wouldn't plan around it. A 24 GB machine is the comfortable size. Google also makes smaller Gemma 4 sizes (E2B and E4B) for small devices. I haven't measured those.

How much RAM does Gemma 4 26B need on a Mac?

Outlier lists 16 GB as the minimum for the 4-bit 26B-A4B. The tier catalog records a measured peak of about 16 GB while it answers, so 24 GB is where there is real headroom for macOS and your other apps.

How fast is Gemma 4 on Apple Silicon?

43.6 tokens a second for the 26B-A4B on an M1 Ultra, the median of 20 runs on 2026-08-24. That is about twice the speed of a 27B dense model on the same Mac. Other chips will differ, and I only measured the M1 Ultra.

Is Gemma 4 good for coding?

For short functions, yes: 21 of 23 in my set. For fixing a bug in a real repository, no: 0 of 30. Chat coding and agentic coding turned out to be different skills.

Is Gemma 4 free to run locally?

The weights are free to download, and the mlx-community repo is tagged Apache 2.0. The MLX route above costs nothing. In Outlier, Quick is a Pro tier, while Nano and Lite are free.

Receipts: Speed for Quick: median of 20 runs on a Mac Studio with an M1 Ultra, through the shipped Outlier 1.11.804 bundle on 2026-08-24, routed by model id and confirmed in the engine log. Speed for Nano, Lite and Core: my runs from the 2026-06-09 batch on the same Mac, 4-bit MLX, one request at a time. File sizes are read from the shipping tier catalog (Outlier 1.11.909). The 16 GB peak is the measured generation footprint recorded in the tier catalog. MMLU: n=300, stratified across 50 of 57 subjects, April 2026, answers read from chat replies, 0 errors, Wilson 95% interval. Short code: n=23, my own problem set, M1 Ultra, shipping build, August 2026. Repository bugs: SWE-bench Verified, n=30, blind (no test patch), official Docker harness, August 2026. HumanEval: n=164, April 2026, fenced-code extraction. The roughly 3 seconds of thinking is 146 tokens divided by 43.6 tokens a second. The Hugging Face repo name and licence tag were read on 2026-10-10. Every figure comes from the same Mac Studio and one chip. None of it is a rate for your Mac.

Try Outlier free

Free: Nano + Lite. Pro: $249 once · or 4 × $62.25. macOS 26+. In the US, Klarna or Afterpay at checkout: four payments, two weeks apart. Refund window: 30 days, no questions asked.

Download free (290 MB) Buy Pro

New models worth running