Fastest published Mac-family result
Mac Studio M3 Ultra 96GB maps to 16.7 tok/s from the strongest published chip-family row for GLM-5.
GLM-5 ranked across the Mac lineup at the best practical quantization, using the best available runtime evidence. Historical baseline selected; model picker is focused on current-market choices.
Historical baseline selected: GLM-5. Default model choices remain current-market; other historical models stay hidden.
| Rank | Mac | Score | Quant | Tok/s | Runtime | Fits | Headroom | Context | Evidence | Price | Why it ranks here |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Mac Studio M3 Ultra 256GB | 120 | IQ2_XS | 13.2 tok/s Fastest evidence path: IQ2_XS · 13.2 tok/s · MLX · Estimated | MLX | Fits | 43.3 GB | 9k | Estimated | $7,499 | IQ2_XS is the current best practical quantization. 13.2 tok/s is estimated from nearby benchmark coverage. 43.3 GB headroom remains at this quantization. |
| 2 | Mac Mini M4 16GB | 0 | F32 | — | MLX | No | -2795.1 GB | — | Estimated | $499 | GLM-5 does not fit on Mac Mini M4 16GB at the current practical quantization. |
| 3 | Mac Mini M4 24GB | 0 | F32 | — | MLX | No | -2787.1 GB | — | Estimated | $599 | GLM-5 does not fit on Mac Mini M4 24GB at the current practical quantization. |
| 4 | Mac Mini M4 32GB | 0 | F32 | — | MLX | No | -2779.1 GB | — | Estimated | $799 | GLM-5 does not fit on Mac Mini M4 32GB at the current practical quantization. |
| 5 | MacBook Air M4 16GB 13-inch | 0 | F32 | — | MLX | No | -2795.1 GB | — | Estimated | $1,099 | GLM-5 does not fit on MacBook Air M4 16GB 13-inch at the current practical quantization. |
| 6 | MacBook Air M4 24GB 13-inch | 0 | F32 | — | MLX | No | -2787.1 GB | — | Estimated | $1,299 | GLM-5 does not fit on MacBook Air M4 24GB 13-inch at the current practical quantization. |
| 7 | MacBook Air M4 16GB 15-inch | 0 | F32 | — | MLX | No | -2795.1 GB | — | Estimated | $1,299 | GLM-5 does not fit on MacBook Air M4 16GB 15-inch at the current practical quantization. |
| 8 | Mac Mini M4 Pro 24GB | 0 | F32 | — | MLX | No | -2787.1 GB | — | Estimated | $1,399 | GLM-5 does not fit on Mac Mini M4 Pro 24GB at the current practical quantization. |
| 9 | MacBook Air M4 32GB 13-inch | 0 | F32 | — | MLX | No | -2779.1 GB | — | Estimated | $1,499 | GLM-5 does not fit on MacBook Air M4 32GB 13-inch at the current practical quantization. |
| 10 | MacBook Air M4 24GB 15-inch | 0 | F32 | — | MLX | No | -2787.1 GB | — | Estimated | $1,499 | GLM-5 does not fit on MacBook Air M4 24GB 15-inch at the current practical quantization. |
| 11 | Mac Mini M4 Pro 48GB | 0 | F32 | — | MLX | No | -2763.1 GB | — | Estimated | $1,599 | GLM-5 does not fit on Mac Mini M4 Pro 48GB at the current practical quantization. |
| 12 | MacBook Air M4 32GB 15-inch | 0 | F32 | — | MLX | No | -2779.1 GB | — | Estimated | $1,699 | GLM-5 does not fit on MacBook Air M4 32GB 15-inch at the current practical quantization. |
| 13 | MacBook Pro M4 Pro 24GB 14-inch | 0 | F32 | — | MLX | No | -2787.1 GB | — | Estimated | $1,999 | GLM-5 does not fit on MacBook Pro M4 Pro 24GB 14-inch at the current practical quantization. |
| 14 | Mac Studio M4 Max 36GB | 0 | F32 | — | MLX | No | -2775.1 GB | — | Estimated | $1,999 | GLM-5 does not fit on Mac Studio M4 Max 36GB at the current practical quantization. |
| 15 | MacBook Pro M4 Pro 48GB 14-inch | 0 | F32 | — | MLX | No | -2763.1 GB | — | Estimated | $2,499 | GLM-5 does not fit on MacBook Pro M4 Pro 48GB 14-inch at the current practical quantization. |
| 16 | MacBook Pro M4 Pro 24GB 16-inch | 0 | F32 | — | MLX | No | -2787.1 GB | — | Estimated | $2,499 | GLM-5 does not fit on MacBook Pro M4 Pro 24GB 16-inch at the current practical quantization. |
| 17 | Mac Studio M4 Max 48GB | 0 | F32 | — | MLX | No | -2763.1 GB | — | Estimated | $2,499 | GLM-5 does not fit on Mac Studio M4 Max 48GB at the current practical quantization. |
| 18 | MacBook Pro M4 Max 36GB 14-inch | 0 | F32 | — | MLX | No | -2775.1 GB | — | Estimated | $2,999 | GLM-5 does not fit on MacBook Pro M4 Max 36GB 14-inch at the current practical quantization. |
| 19 | MacBook Pro M4 Pro 48GB 16-inch | 0 | F32 | — | MLX | No | -2763.1 GB | — | Estimated | $2,999 | GLM-5 does not fit on MacBook Pro M4 Pro 48GB 16-inch at the current practical quantization. |
| 20 | Mac Studio M4 Max 64GB | 0 | F32 | — | MLX | No | -2747.1 GB | — | Estimated | $2,999 | GLM-5 does not fit on Mac Studio M4 Max 64GB at the current practical quantization. |
| 21 | MacBook Pro M4 Max 48GB 14-inch | 0 | F32 | — | MLX | No | -2763.1 GB | — | Estimated | $3,499 | GLM-5 does not fit on MacBook Pro M4 Max 48GB 14-inch at the current practical quantization. |
| 22 | MacBook Pro M4 Max 36GB 16-inch | 0 | F32 | — | MLX | No | -2775.1 GB | — | Estimated | $3,499 | GLM-5 does not fit on MacBook Pro M4 Max 36GB 16-inch at the current practical quantization. |
| 23 | MacBook Pro M4 Max 48GB 16-inch | 0 | F32 | — | MLX | No | -2763.1 GB | — | Estimated | $3,999 | GLM-5 does not fit on MacBook Pro M4 Max 48GB 16-inch at the current practical quantization. |
| 24 | Mac Studio M3 Ultra 96GB | 0 | F32 | — | MLX | No | -2715.1 GB | — | Estimated | $3,999 | GLM-5 does not fit on Mac Studio M3 Ultra 96GB at the current practical quantization. |
| 25 | MacBook Pro M4 Max 64GB 16-inch | 0 | F32 | — | MLX | No | -2747.1 GB | — | Estimated | $4,499 | GLM-5 does not fit on MacBook Pro M4 Max 64GB 16-inch at the current practical quantization. |
| 26 | Mac Studio M4 Max 128GB | 0 | F32 | — | MLX | No | -2683.1 GB | — | Estimated | $4,499 | GLM-5 does not fit on Mac Studio M4 Max 128GB at the current practical quantization. |
| 27 | MacBook Pro M5 Max 128GB 16-inch | 0 | F32 | — | MLX | No | -2683.1 GB | — | Estimated | $5,399 | GLM-5 does not fit on MacBook Pro M5 Max 128GB 16-inch at the current practical quantization. |
| 28 | MacBook Pro M4 Max 128GB 16-inch | 0 | F32 | — | MLX | No | -2683.1 GB | — | Estimated | $5,999 | GLM-5 does not fit on MacBook Pro M4 Max 128GB 16-inch at the current practical quantization. |
| 29 | Mac Pro M2 Ultra 192GB | 0 | F32 | — | MLX | No | -2619.1 GB | — | Estimated | $6,999 | GLM-5 does not fit on Mac Pro M2 Ultra 192GB at the current practical quantization. |
Start with the ranked Mac table above, then audit tokens per second, RAM fit, quantization, runtimes, and source links before trusting the model on your Mac.
Quantizations observed: 4bit
Quick take
Fastest published result is 16.7 tok/s on M3 Ultra (512 GB) at 4bit. Smallest published fit is 391.8 GB on M3 Ultra (512 GB). Longest published context on this page is 33k. Published runtimes include MLX. Start with Rankings for the decision, then use the raw rows below to audit the evidence.
Based on 5 external benchmarks; no lab runs yet.
Published runtimes: MLX.
Best Mac shortlist
These are ranked by the fastest published GLM-5 result available for each Mac's chip family. Use them as the search answer, then open the machine page before buying.
Mac Studio M3 Ultra 96GB maps to 16.7 tok/s from the strongest published chip-family row for GLM-5.
Mac Studio M3 Ultra 256GB maps to 16.7 tok/s from the strongest published chip-family row for GLM-5.
Model search answers
GLM-5 currently has a fastest published Mac result of 16.7 tok/s on M3 Ultra (512 GB) at 4bit. Check the raw rows on this page before comparing that number to a different runtime, quantization, or context length.
Mac Studio M3 Ultra 96GB is the fastest published Mac-family answer for GLM-5 on this page right now, at 16.7 tok/s from its chip family. Treat that as a published benchmark starting point, not a universal buying recommendation.
GLM-5 has a smallest published fit of 391.8 GB on M3 Ultra (512 GB). More memory may still be required for longer context, different quantization, or parallel workloads.
Catalog record
This is a reference-only model record. It remains useful for historical benchmarks, migration checks, and audit context, but it is excluded from current frontier packs.
Official model cards tell you what the model is for and which software stacks it targets. Field reality below shows how much Apple Silicon evidence we have so far.
Official brief
We are launching GLM-5, targeting complex systems engineering and long-horizon agentic tasks. Scaling is still one of the most important ways to improve the intelligence efficiency of Artificial General Intelligence (AGI).
Official source · Raw model card
Runtime support mentioned
Official specs
Official takeaways
Official model cards describe intent, capabilities, and supported stacks. They do not prove Apple Silicon speed by themselves.
Field reality on Apple Silicon
GLM-5: 6 Apple Silicon field reports; best reported generation ~20 tok/s; best reported prompt processing ~187 tok/s; reported RAM use ~391.82-415.41GB; seen on M3 ULTRA, Mac Studio M3 ULTRA 512GB; via oMLX.
What practitioners keep saying
Apple Silicon field sources
An accessible M3 Ultra 512GB oMLX report shows GLM-5 running on Apple Silicon with slow single-request latency but materially better throughput under continuous batching and persistent KV cache.
A LocalLLaMA operator reports Unsloth GLM-5 low-bit quants running on M3 Ultra, including Q2 around 20 tok/s, while the source leaves key setup details unspecified.
GLM-5 is no longer theoretical on Apple Silicon; operators are already running it on M3 Ultra-class desktops and comparing the experience against frontier hosted models.
Runtime/source notes to verify
The official GLM-5 launch post says the model weights are available for local deployment and names vLLM and SGLang as supported serving frameworks, but it does not establish Mac throughput or fit.
Runtime mentions in the field
Hardware mentioned in reports
What would improve confidence
Current published coverage
Published chip coverage includes M3 Ultra (512 GB). Fastest published row is 16.7 tok/s on M3 Ultra (512 GB) at 4bit. Lowest published RAM requirement is 391.8 GB on M3 Ultra (512 GB). Catalog context window is 33k.
Rows stay below the ranking because this page is answer-first. Use them to inspect exact chips, quantizations, runtimes, and sources.
| Chip | Quant | Avg tok/s | Runtime | Source |
|---|---|---|---|---|
| M3 Ultra (512 GB) | 4bit | 16.7 tok/s | MLX | ref |
| M3 Ultra (512 GB) | 4bit | 13.7 tok/s | MLX | ref |
| M3 Ultra (512 GB) | 4bit | 13.2 tok/s | MLX | ref |
| M3 Ultra (512 GB) | 4bit | 12.0 tok/s | MLX | ref |
| M3 Ultra (512 GB) | 4bit | 10.7 tok/s | MLX | ref |
Chips with published results for GLM-5
Data
benchmarks.json — full dataset · models.json — model summaries · benchmarks.csv — CSV export