Qwen3.8-27B at 4-bit MLX and Meta Muse Glimmer 30B are the picks for a 24GB Apple Silicon Mac. Devstral Small 2 was retired in March 2026.
Our practical pick for a 24GB Apple Silicon Mac is Qwen3.8-27B in 4-bit MLX, at roughly 16.1GB of weights, with Meta Muse Glimmer 30B as a co-primary alternative at 17GB in NVFP4. Both shipped in August 2026, both are Apache 2.0, and both leave enough unified memory free for a working context window and the rest of the system.
This page first ran on 17 May 2026, when the pick was Qwen3.6-27B. That model still works and every benchmark figure we published for it still checks out, but it was superseded on 14 August 2026. Dates appear throughout this article on purpose: in this category a recommendation without a date is a recommendation you cannot audit.
Methodology: what is verified, what is estimated
Memory figures in this article are file sizes and download sizes published by the model vendors and by the Ollama, LM Studio and MLX communities, checked on 7 October 2026. Resident memory during inference runs higher than the file size once the KV cache and runner overhead are counted. Speed figures are only quoted where a vendor or a named tester measured them on stated hardware.
The ranking prioritizes four factors:
- Verified model specifications, including parameter count, license and context window
- Model quality signals, especially coding, reasoning and agentic workflow benchmarks
- Practical local fit, including quantized weight size and KV-cache overhead
- Mac usability, including support through local runners such as MLX, Ollama, LM Studio or llama.cpp
Exact tokens per second will vary sharply by chip generation, runner, quantization, context length, thermal state and background workload. Small 3B-4B models should feel much faster than 27B-30B models, but any precise speed claim should be treated as a test result only if measured on the same hardware and software stack.
Why 24GB is tighter than it sounds
Apple Silicon uses unified memory. The CPU, GPU, Neural Engine, operating system and applications all draw from the same memory pool. That makes local AI convenient, because the GPU can access large shared memory without a separate VRAM limit, but it also means the full 24GB is never available solely for model inference.
The second constraint is KV cache, the memory used to store attention state as context grows. KV-cache usage increases with context length, layer count, hidden size, number of attention heads and cache precision. A model may advertise a 262,144-token context window, but using anywhere near that limit locally can consume far more memory than the model weights alone.
One more note on the hardware itself. As of Apple's 25 August 2026 announcements, shipping 22 September 2026, 24GB is no longer the default Mac configuration. The M6 Mac mini starts at 16GB with 24GB as a paid upgrade, and 24GB is the floor on the M5 Pro Mac mini, which also carries substantially more memory bandwidth. For that lower tier, see the best local LLM picks for 16GB VRAM. The M5 and M6 generations add per-core GPU Neural Accelerators, which make this tier meaningfully faster than it was in May 2026. If you are reading this while choosing a machine, 24GB is a deliberate configuration choice now, not something you get by default.
Shortlist: the most relevant models for 24GB Apple Silicon
| Model | Verified model facts | Local fit on 24GB Mac | Best role |
|---|---|---|---|
| Qwen3.8-27B | Released 14 August 2026, dense 27B, Apache 2.0, 262,144-token native context extensible to 1M, native vision-language | Comfortable at 4-bit MLX (16.1GB) or Q4_K_M (16.8GB plus a 0.93GB vision projector) | Best overall practical pick |
| Meta Muse Glimmer 30B | Released 10 August 2026, roughly 29.6B dense plus a separate 1.8B ViT-G/14 perception encoder, Apache 2.0, 131,072-token context, text and image in | Comfortable at NVFP4 (17GB) or Q4_K_M (16.8GB plus 1.4GB for the vision encoder) | Co-primary pick, strongest agentic option |
| Gemma 4 26B A4B | Released 31 March 2026, 25.2B total Mixture-of-Experts with 3.8B active parameters, Apache 2.0, 256K context | Comfortable at 4-bit MLX (15.3GB) or QAT Q4 (14.2GB); fastest of the large picks | Best document and RAG model |
| Qwen3.6-27B | Released 22 April 2026, roughly 27.8B dense, 55.6GB in BF16, Apache 2.0, 262,144-token context | Still fits at Q4_K_M (16.8GB), but superseded | Previous generation, keep if already installed |
| Gemma 4 12B Unified | Released 3 June 2026, roughly 11.95B, Apache 2.0, 256K context, encoder-free with native audio and video input | Very comfortable at Google QAT Q4 (6.98GB) or 4-bit MLX (6.74GB) | Fast companion, only local pick with native audio |
| Devstral Small 2 24B | Released 9 December 2025, 24B, Apache 2.0, 256K context, 68.0% SWE-bench Verified. Retired by Mistral on 31 March 2026 | Weights still run, but Mistral's own card specifies a 32GB Mac | Legacy coding model, no longer recommended |
| IBM Granite 4.2 30B | Published August 2026, 30B dense (29.3B parameters in safetensors), Apache 2.0, 128K native context (512K with long-context extension) | Fits: IBM's own 4-bit MLX build is about 16.5GB, Q4_K_M GGUF 17.7GB | Also fits; no benchmark data here, so we do not rank it above the picks |
| DeepSeek V4.1-Flash | Released 10 September 2026, 552B backbone, 8B active in prefill and 16B in decode, MIT, 1M context | Not practical for 24GB local use | Server/API model, not a Mac default |
Large MoE systems may activate only a small portion of their parameters per token, but the machine still needs access to the full weight set. Active-parameter count is not the same thing as local memory footprint.
Why Qwen3.8-27B is the practical pick
Qwen3.8-27B lands in the zone that matters for a 24GB Mac: large enough to be meaningfully stronger than small local assistants, but small enough that a 4-bit build leaves real working room.
Start with the footprint, not the full-precision size. The mlx-community/Qwen3.8-27B-4bit build is 16.1GB, and the Q4_K_M GGUF is 16.8GB plus a 0.93GB vision projector. Ollama's qwen3.8:27b-nvfp4 and qwen3.8:27b-q4_K_M tags are both 18GB packaged. These are not fringe community conversions any more: mlx-community, lmstudio-community and the Ollama library all carry official-ecosystem Apple Silicon builds, and Ollama shipped optimized MLX variants on the day the weights went public.
The model's strongest case is coding and agentic work. Its published results include 61.7 on SWE-bench Pro, 73.0 on Terminal-Bench 2.1, 90.3 on LiveCodeBench v6 and 89.2 on GPQA Diamond, alongside 84.3 on OSWorld-Verified and 94.6 on MathVision for multimodal tasks. Qwen has not published a SWE-bench Verified figure for the 3.8 generation, so we are not quoting one.
For comparison, the Qwen3.6-27B numbers we published in May still verify exactly against its model card: 77.2% on SWE-bench Verified, 53.5% on SWE-bench Pro, 59.3 on Terminal-Bench 2.0, 83.9 on LiveCodeBench v6, 86.2 on MMLU-Pro, 87.8 on GPQA Diamond and 94.1 on AIME 2026. Where the two generations overlap, 3.8 wins every time. Artificial Analysis scores Qwen3.8 27B at 34 on its Intelligence Index against 22 for Qwen3.6 27B.
Two practical notes. Qwen3.8 is a native vision-language model, so it reads images and video, which Qwen3.6 did not. And there is no smaller Qwen3.8 sibling — no 4B, 8B or 14B in this generation. If you want something lighter from Qwen you have to drop back to a Qwen3.6 or Qwen3.5 checkpoint.
Meta Muse Glimmer 30B is the co-primary pick
Muse Glimmer 30B, released 10 August 2026, is the first open model from Meta Superintelligence Labs and the strongest new entrant in this class. It is a dense model of roughly 29.6B parameters plus a separate 1.8B ViT-G/14 perception encoder, released under Apache 2.0 with a 131,072-token context, taking text and images and producing text.
It earns the co-primary slot on a detail no other vendor in this list offers: Meta publishes its own quantized weights and explicitly labels one of them for 24GB devices, naming MacBook M4-Max and M5-Max machines as targets. The Muse-Glimmer-30B-KQuant-17GB-Q4_K_M.gguf build is 16.8GB, and Ollama's muse-glimmer:30b-nvfp4 tag is 17GB. Budget a further 1.4GB for the vision projector if you want image input.
Muse Glimmer also ships a DFlash speculative-decoding drafter, a block-diffusion model that proposes 16 tokens per forward pass. Meta reports 1.5x on an Apple M4 Max and 1.8x on an M5 Max at identical output quality. The drafter costs memory — Ollama's muse-glimmer:30b-q4_K_M-dflash tag is 20GB against 18GB without — so on 24GB it is a trade against context length rather than a free speedup.
Pick Muse Glimmer if your work is agentic and local-first. Pick Qwen3.8-27B if you want the longer context window and the stronger published coding numbers.
Gemma 4 26B A4B is the strongest document model
Gemma 4 26B A4B is the best choice when the workload is less about coding agents and more about RAG, document Q&A, multimodal understanding and long-form analysis.
Released 31 March 2026, it uses a Mixture-of-Experts design with 25.2B total and 3.8B active parameters during inference, an Apache 2.0 license, and a 256K context window of its own — not inherited from a larger sibling. On a bandwidth-bound Mac that 3.8B active count is the important number: it decodes far faster than a dense 27B while occupying similar memory.
Two things changed since May. Google shipped quantization-aware training checkpoints for the entire Gemma 4 family in June 2026, which give better quality at the same footprint than the community post-training quants most guides still point at. The QAT Q4 build of 26B A4B is 14.2GB, the lightest of the large picks here. And Gemma 4 12B Unified arrived on 3 June 2026: an encoder-free model of roughly 11.95B parameters with native audio and video input, at 6.98GB in Google's QAT Q4 GGUF. Nothing else in this article takes audio.
Devstral Small 2 was retired, and there is no replacement at this size
Our May recommendation hedged on Devstral Small 2 by telling readers to verify the served endpoint. That hedge was too soft, and we are correcting it.
Mistral deprecated Devstral Small 2 on 27 February 2026 and retired it on 31 March 2026 — seven weeks before this article first published. Mistral's own model list names Mistral Medium 3.5, released 30 April 2026, as the replacement. Medium 3.5 is a 128B dense model and does not fit a 24GB Mac. Neither does Mistral Small 4, which is a 119B Mixture-of-Experts model.
The weights are not gone. mistralai/Devstral-Small-2-24B-Instruct-2512 is still downloadable under Apache 2.0, still scores 68.0% on SWE-bench Verified, and will still run locally — but Mistral's own model card specifies a single RTX 4090 or a Mac with 32GB RAM, which was always the honest reading of its local fit. A retired hosted endpoint does not stop open weights from working, but it does mean no further updates, no vendor support and no successor at this size.
Mistral has shipped no local-scale coding model since. Everything in its release feed between May and September 2026 is either much larger or a different category entirely. If you came here for a local coding model, use Qwen3.8-27B or Muse Glimmer 30B instead.
The fast companion slot
Phi-4-mini-instruct remains a genuinely good small model — 3.8B parameters, 128K context, MIT licensing, all still accurate — but it dates to February 2025 and it is no longer the obvious default. There is no Phi-5.
Better options in the same slot now:
- Gemma 4 12B Unified, at 6.98GB in Google's QAT Q4 build, is the strongest of these and the only one that takes audio
- MiMo-V2.6-Distill-Qwen-9B (Xiaomi, MIT, 21 September 2026) is a 9.4B model made by supervised fine-tuning of Qwen3.5-9B; it is 18.8GB in BF16, so it fits 24GB, but Xiaomi describes it as a starting point for open research in agentic reinforcement learning. Treat it as a research checkpoint, not a pick
- Gemma 4 E4B, at 4.22GB in the QAT build, is the speed choice on constrained machines
- LFM2.5-2.6B, released 4 August 2026, ships a DSpark drafter with measured numbers on an M4 Max MacBook Pro: 61 to 139 tokens per second. Note its LFM Open License v1.0 is not Apache 2.0
Use the small model for command generation, short summaries, lightweight RAG, note cleanup, simple scripts and private chat. Keep Qwen3.8-27B or Muse Glimmer 30B for harder reasoning, deeper coding and document-heavy workflows.
A two-model setup makes more sense than forcing one large model to handle everything. Run the small model for latency. Run the larger model when quality matters.
Quantization: the Q4/Q5/Q6 ladder is no longer the whole menu
In May this section said Q4_K_M with Q5 as an upgrade. That advice is now incomplete.
Ollama's MLX engine update on 11 June 2026 introduced NVFP4 on Apple Silicon, and the measured claims are specific: NVFP4 roughly halves the quality loss of 4-bit quantization relative to unquantized BF16, measured by perplexity on Gemma 4 12B, and generates about 20% faster than Q4_K_M on the updated engine, with several operations fused into single Metal kernels by MLX's just-in-time compiler. MXFP4 and MXFP8 are also available as tags, though MXFP8 builds run 32GB and up and are a 48GB-tier option.
Current guidance for 24GB:
- NVFP4 where the model publishes it — Qwen3.8-27B and Muse Glimmer 30B both do, and on Muse Glimmer it is smaller than Q4_K_M (17GB against 18GB packaged)
- 4-bit MLX as the Apple-native default, and the smallest of the Qwen3.8 builds at 16.1GB
- Q4_K_M as the portable fallback that works everywhere
- QAT Q4 for Gemma 4, which beats generic post-training quantization at the same size
- Q5 and above are the wrong trade on 24GB for a 27B-30B model; they eat the headroom the KV cache needs
On context, we are raising our guidance. In May we said 8K-16K. With a 4-bit MLX build occupying roughly 16GB, 16K-32K is a realistic working range, and testers have reported Qwen3.8-27B running at 64K context on a 24GB M4 Pro at around 12 tokens per second using Q4_K_M with flash attention enabled. The levers that make this possible are KV-cache quantization and flash attention; turn both on before you lower the context window. The full 262,144-token window remains out of reach — KV cache alone runs to roughly 16GB at that length. Our VRAM calculator for open-weight models across quantization formats can size this for other context lengths.
What not to run locally on 24GB
DeepSeek is an important open-weight family, but not a realistic 24GB Mac recommendation. DeepSeek-V4.1-Flash, released 10 September 2026 under MIT, has a 552B-parameter backbone, activates 8B parameters in prefill and 16B in decode, and supports 1M-token context, per its model card. The larger V4-Pro is 1.6T parameters, and the original V4-Flash (24 April 2026) was 284B total with 13B active. All are in a very different hardware category.
The weeks since the last update brought a wave of large open-weight releases. None of them is a 24GB Mac model, so none changes the pick (checked 7 October 2026):
- GLM-5.3-Flash (Z.ai, MIT, weights published in August 2026) has about 320B total and 18B active parameters, and Hugging Face lists roughly 321B parameters in its checkpoint. Active parameters do not shrink the download; the weights are hundreds of gigabytes.
- MiMo-V2.6-Flash (Xiaomi, MIT, published 21 September 2026 as the
MiMo-V2.6-Flash-RLand-MOPDcheckpoints) is a sparse MoE that the model card lists as 309B total and 15B active with 1M-token context. Hugging Face shows about 311B parameters and roughly 178GB of storage. Xiaomi's larger MiMo-V2.6-Pro is bigger still (see our coverage). - MiniMax M3, released in June 2026 under MiniMax's community licence, is about 428B total and 23B active. We found no MiniMax M3.1 on MiniMax's Hugging Face organisation or in the M3 model card, so we do not cover one.
- Qwen3.8-Flash-Next (Qwen, released 24 August 2026 under the qwen-community-1.0 licence, not Apache 2.0) is 125B with 6B active, plus a 51B n-gram embedding and 4B MTP; the checkpoint is about 180GB. It is the only other Qwen3.8 release besides the 27B, so there is still no smaller sibling.
- Ling 3.1 Flash (Ant Group, announced 30 September 2026) is a roughly 560B MoE with about 25B active per token. Ant says it will open-source the model after a free trial period; we found no published weights, so it is not an open-weight option yet. See our three-way comparison with GLM-5.3-Flash and Qwen3.8-Flash-Next. Its predecessor, Ling-3.0-flash, is open under MIT but is a 124B MoE, also far beyond 24GB.
Two more that look tempting and are not:
- Poolside Laguna XS 2.1, released 2 July 2026, is a genuinely strong agentic coding model — 33B total with 3B active, OpenMDW-1.1 license, 262,144-token context. But its recommended Q4_K_M build is 20.3GB, and Poolside itself specifies a 36GB Mac. On 24GB that is a knife-edge once the OS and KV cache are counted.
- Gemma 4 31B, released 31 March 2026, fits at 17.3GB in its QAT build, but as a dense 30.7B model it decodes slower than the 26B A4B at a larger footprint. Prefer the MoE.
The relevant question is not whether a model can theoretically be downloaded. It is whether it can run with enough context, speed and stability to be useful while the rest of the system remains usable.
Recommended setup
| Use case | Pick |
|---|---|
| Best overall local assistant | Qwen3.8-27B, 4-bit MLX (16.1GB) |
| Best agentic and local-first alternative | Meta Muse Glimmer 30B, NVFP4 (17GB) |
| Best coding-focused model | Qwen3.8-27B or Muse Glimmer 30B — Devstral Small 2 was retired on 31 March 2026 |
| Best RAG/document model | Gemma 4 26B A4B, QAT Q4 (14.2GB) |
| Only local pick with native audio input | Gemma 4 12B Unified, QAT Q4 (6.98GB) |
| Fastest everyday assistant | Gemma 4 E4B (4.22GB) or LFM2.5-2.6B |
| Best speed-per-GB on a bandwidth-bound Mac | A sparse MoE — Gemma 4 26B A4B activates 3.8B of 25.2B parameters per token and decodes faster than a dense 27B at similar memory |
| Best quantization when published | NVFP4, then 4-bit MLX, then Q4_K_M |
| Best context setting for 27B-30B models | 16K-32K, with KV-cache quantization and flash attention on |
| Best convenience runners | LM Studio or Ollama |
| Best control-oriented runner | llama.cpp |
| Apple Silicon inference engine | MLX — it has been Ollama's Apple Silicon backend since v0.19 on 30 March 2026, and LM Studio ships it alongside llama.cpp. Use MLX-LM directly when you want maximum control |
On runners specifically: our May advice treated MLX as a future option to adopt once builds matured. That framing is out of date. Ollama replaced its llama.cpp Metal backend with MLX in v0.19 on 30 March 2026, reporting a 93% improvement in decode and 57% in prefill, and LM Studio ships both engines. If you are running Ollama or LM Studio on an Apple Silicon Mac today, you are already running MLX. For the deeper tradeoffs between the two runtimes, see our MLX vs llama.cpp benchmarks on Apple Silicon. As of 7 October 2026 the latest releases we checked are Ollama v0.40.0 (25 September) and MLX v0.32.3 (29 September); llama.cpp ships continuous builds, so check its releases page. LM Studio 0.4.22, released 28 August 2026, added support for DFlash, DSpark and MTP drafter models, which is what makes the speculative-decoding numbers above reachable from a GUI.
Worth knowing for macOS 27: at WWDC in June 2026 Apple opened its Foundation Models framework to third-party backends and is open-sourcing an MLX language model adapter, which lets a native app call any mlx-community model through Apple's own API rather than a separate runner.
Final verdict
For a 24GB Apple Silicon Mac as of 7 October 2026, the most defensible recommendation is Qwen3.8-27B in 4-bit MLX, with Meta Muse Glimmer 30B as an equally serious alternative — particularly if your work is agentic, or if you value a vendor that ships a quantization built for your exact memory tier.
Users who mainly work with documents should test Gemma 4 26B A4B in its QAT build. Anyone who needs audio input has exactly one local option, Gemma 4 12B Unified. Anyone who values instant responses should keep a small model such as Gemma 4 E4B installed as a companion.
Do not start a new project on Devstral Small 2. It was retired on 31 March 2026 and Mistral has shipped no local-scale successor.
Update, 7 October 2026: DeepSeek-V4.1-Flash, GLM-5.3-Flash, MiMo-V2.6-Flash, MiniMax M3, Qwen3.8-Flash-Next and the announced Ling 3.1 Flash are all far too large for this tier (see the section on what not to run). One new model does fit, IBM Granite 4.2 30B (about 16.5GB in IBM's 4-bit MLX build), but we have no benchmark data to rank it above the picks, so they stand.
A 24GB Mac can run a capable private AI assistant in 2026, and the August model releases made it a better machine than it was in May. Readers with more headroom should see the 128GB Apple Silicon tier, where Qwen3.6 leads. But if the goal is serious agentic coding, long debugging loops, large repositories or high-context RAG, 48GB remains the more comfortable tier. The difference is not just whether a model loads. It is whether it stays useful after the prompt, the tools, the files, the browser and the operating system all compete for the same memory.