M1 Ultra
How fast is 70B Q4_K_M on M1 Ultra?
M1 Ultra 70B Q4_K_M / 4-bit
64 GB row: 12.6 tok/s on LM Studio at 4k context with 4bit. Community unverified evidence.
Bench is the Apple Silicon LLM benchmark audit surface for exact questions like M1 Ultra 70B Q4_K_M tokens/s, M2 Max 70B Q4_K_M speed, and Qwen3-Coder-Next on M3 Ultra. Search the measured rows behind Silicon Score: Mac chip, model, quantization, runtime, RAM required, context window, prompt speed, decode tok/s, source, and evidence class.
M1 Ultra
M1 Ultra 70B Q4_K_M / 4-bit
64 GB row: 12.6 tok/s on LM Studio at 4k context with 4bit. Community unverified evidence.
M2 Max
M2 Max 70B Q4_K_M / 4-bit
64 GB row: 8.8 tok/s on LM Studio with 4bit. Community unverified evidence.
M3 Ultra
M3 Ultra 70B 4-bit
512 GB row: 15.5 tok/s on LM Studio at 8k context with 4bit. Community unverified evidence.
M3 Ultra
M3 Ultra Qwen3-Coder-Next 4-bit
256 GB row: 74.0 tok/s on MLX with 4bit. Community unverified evidence.
Catalog state
Benchmark rows
432
Open issues
15
Reviewed issues
15
Research queue
30
Workflow comparisons
4
Model scorecards
10
Featured lab
The featured environment is an operational pointer for Rankings and Bench. It stays on the owned M5 Max until the planned Mac Studio is physically verified and has publishable clean-recording rows.
Current default
active
Public Rankings default and active first-party lab target: macbook-pro-m5-max-128gb-16.
Runbook
data/ops/bitter-mill/current-frontier-m5-max-2026-05.jsonNext command
npm run bench:hygiene:doctor -- --clean-recording --output .local/benchmark-hygiene/current-frontier-m5-max-2026-05.doctor.jsonPlanned successor
planned
Expected June 2026; provisional id mac-studio-m4-ultra-256gb stays out of the public catalog until verified.
Runbook
data/ops/bitter-mill/planned/m4-ultra-256gb-frontier-2026-06.jsonNext command
npm run bench:hygiene:doctor -- --clean-recording --output .local/benchmark-hygiene/m4-ultra-256gb-frontier-2026-06.doctor.jsonPromotion gate
planned
In June 2026 the local lab is expected to receive a 256GB Mac Studio provisionally labeled M4 Ultra, replacing the M5 Max 128GB MacBook Pro as the featured environment only after arrival, hardware detection, and clean first-party evidence justify changing the public default.
Next command
npm run validate:lab-environmentsCapture system_profiler hardware evidence for chip name, model identifier, GPU cores, and 256GB unified memory after the Mac Studio arrives.
Keep the current M5 Max 128GB machine as the featured public default until the Mac Studio has at least one clean-recording first-party Bitter Mill or Silicon Score Lab benchmark row.
Record clean-recording hygiene with memory pressure, thermal/performance warning, local-inference process, final snapshot, and zero swap I/O evidence before any benchmark row is promoted.
Verify Apple or hardware-detection naming before creating or promoting a public data/machines.json record for the planned machine.
Audit now
This board merges unresolved quality issues, frontier candidates without enough evidence, and active operator tasks. It prioritizes the gaps most likely to change Rankings or Bench confidence.
Run Bitter Mill current-frontier batch on the owned M5 Max
Qwen3.6, MiniMax M2.7, Gemma 4 including E2B/E4B, Mistral Small 4, Ministral 3 compact models, Llama 4 Scout, gpt-oss 120B/20B, Nemotron Cascade 2, GLM-4.5-Air, Magistral Small, Devstral Small 2, and current Qwen coding/dense anchors should be reproduced from first-hand Bitter Mill runtime, quantization, and context sweeps before they become first-party measured evidence.
22 models · 1 source
Status: Ready after hygiene · Run the listed clean-recording gate before benchmark execution, discovery, import, or publication changes.
Runbook
data/ops/bitter-mill/current-frontier-m5-max-2026-05.jsonNext command
npm run bench:hygiene:doctor -- --clean-recording --output .local/benchmark-hygiene/current-frontier-m5-max-2026-05.doctor.jsonMonitor command
npm run bench:hygiene:session -- monitor --clean-recording --label bitter-mill-current-frontier-m5-max-2026-05 --sample-interval-ms 15000 --output data/ops/bitter-mill/incoming/current-frontier-m5-max-2026-05/system-hygiene.jsonHygiene gate: Clean recording · clean intake root · memory metrics present · pre-existing swap warning-only · passing final snapshot · passing monitor samples · 0 swap I/O pages · no thermal/performance throttling
Gate: The hygiene doctor reports clean recording readiness before Bitter Mill is opened; if it reports excessive swap allocation, memory pressure, stale local inference processes, or thermal/performance warnings, stop before any inference run. / The all-runbooks readiness audit checks every Bitter Mill runbook with one captured machine and clean-recording system-hygiene snapshot before any runbook is opened.
Run planned M4 Ultra 256GB frontier arrival batch
When the 256GB Mac Studio arrives, run a clean-recording Bitter Mill sweep across the active M5 current-frontier target set including MiniMax M2.7 plus high-memory and fit-boundary extras, runtimes, quantizations including current MLX dynamic low-bit profiles, and context ladders before changing Silicon Score's featured environment.
26 models · 1 source
Status: Blocked until arrival · M4 Ultra 256 GB stays queued until the machine is physically present, hardware identity is verified, and clean-recording preflight passes.
Runbook
data/ops/bitter-mill/planned/m4-ultra-256gb-frontier-2026-06.jsonNext command
npm run bench:hygiene:doctor -- --clean-recording --output .local/benchmark-hygiene/m4-ultra-256gb-frontier-2026-06.doctor.jsonMonitor command
npm run bench:hygiene:session -- monitor --clean-recording --label bitter-mill-m4-ultra-256gb-frontier-2026-06 --sample-interval-ms 15000 --output data/ops/bitter-mill/incoming/m4-ultra-256gb-frontier-2026-06/system-hygiene.jsonHygiene gate: Clean recording · clean intake root · memory metrics present · pre-existing swap warning-only · passing final snapshot · passing monitor samples · 0 swap I/O pages · no thermal/performance throttling
Gate: The hygiene doctor reports clean recording readiness before Bitter Mill is opened on the arrived Mac Studio. / The active all-runbooks readiness audit passes for checked-in active runbooks, then the planned M4 Ultra runbook's explicit check passes on the arrived Mac Studio.
Prepare M4 Ultra 256GB Mac Studio as the next featured lab environment
In June 2026 the local lab is expected to receive a 256GB Mac Studio provisionally labeled M4 Ultra, replacing the M5 Max 128GB MacBook Pro as the featured environment only after arrival, hardware detection, and clean first-party evidence justify changing the public default.
1 source
Status: Blocked until arrival · M4 Ultra 256 GB stays queued until the machine is physically present, hardware identity is verified, and clean-recording preflight passes.
Next command
npm run validate:lab-environmentsHygiene gate: Clean recording · pre-existing swap warning-only · passing final snapshot · 0 swap I/O pages · no thermal/performance throttling
Gate: data/ops/lab-environments.json keeps the M5 Max 128GB MacBook Pro as current active featured environment until the M4 Ultra 256GB Mac Studio is physically present. / Arrival capture records system_profiler chip name, GPU cores, model identifier, and 256GB unified memory before any featured_local_benchmark or FEATURED_MACHINE_ID change.
Upgrade Qwen3.5-122B-A10B reference rows with Bitter Mill runs
Qwen3.5-122B-A10B is a high-memory frontier MoE candidate with trusted-reference rows and field speed reports, and this queued Bitter Mill import will test whether it belongs in high-end Mac advice.
1 model · 1 source
Status: Ready after hygiene · Run the listed clean-recording gate before benchmark execution, discovery, import, or publication changes.
Next command
npm run bench:hygiene:doctor -- --clean-recording --output .local/benchmark-hygiene/import-bitter-mill-qwen-3-5-122b-a10b.doctor.jsonMonitor command
npm run bench:hygiene:session -- monitor --clean-recording --label import-bitter-mill-qwen-3-5-122b-a10b --sample-interval-ms 15000 --output data/ops/bitter-mill/incoming/qwen-3-5-122b-a10b.system-hygiene.jsonHygiene gate: Clean recording · memory metrics present · pre-existing swap warning-only · passing final snapshot · passing monitor samples · 0 swap I/O pages · no thermal/performance throttling
Gate: The hygiene doctor reports clean recording readiness before Bitter Mill is opened; if it reports excessive swap allocation, memory pressure, stale local inference processes, or thermal/performance warnings, stop before any inference run. / The hygiene monitor starts before Bitter Mill is opened and stops only after the model export lands; a same-base monitor hygiene sidecar exists beside the Bitter Mill export.
Upgrade Qwen3.5-397B-A17B reference rows with Bitter Mill runs
Qwen3.5-397B-A17B is the current high-memory Qwen3.5 flagship with sparse Apple Silicon reference evidence, and this queued Bitter Mill import keeps it behind clean first-party reproduction.
1 model · 1 source
Status: Ready after hygiene · Run the listed clean-recording gate before benchmark execution, discovery, import, or publication changes.
Next command
npm run bench:hygiene:doctor -- --clean-recording --output .local/benchmark-hygiene/import-bitter-mill-qwen-3-5-397b-a17b.doctor.jsonMonitor command
npm run bench:hygiene:session -- monitor --clean-recording --label import-bitter-mill-qwen-3-5-397b-a17b --sample-interval-ms 15000 --output data/ops/bitter-mill/incoming/qwen-3-5-397b-a17b.system-hygiene.jsonHygiene gate: Clean recording · memory metrics present · pre-existing swap warning-only · passing final snapshot · passing monitor samples · 0 swap I/O pages · no thermal/performance throttling
Gate: The hygiene doctor reports clean recording readiness before Bitter Mill is opened; if it reports excessive swap allocation, memory pressure, stale local inference processes, or thermal/performance warnings, stop before any inference run. / The hygiene monitor starts before Bitter Mill is opened and stops only after the model export lands; a same-base monitor hygiene sidecar exists beside the Bitter Mill export.
Run Bitter Mill current-reference batch on the owned M5 Max
Qwen 3 32B, Qwen 3 235B-A22B, Devstral Small 1.1, Qwen 3 8B, Nemotron-3-Nano, GLM-4.7-Flash, Phi-4 14B, Mistral Small 3.1, DeepSeek R1 distills, and Qwen 2.5 72B have Apple Silicon evidence or practitioner signals but should be treated as reference/baseline reproduction work rather than the current frontier lane. Qwen3.5-35B-A3B stays in the current-frontier M5 batch as the practical Qwen3.6 displacement comparator.
12 models · 1 source
Status: Ready after hygiene · Run the listed clean-recording gate before benchmark execution, discovery, import, or publication changes.
Runbook
data/ops/bitter-mill/current-reference-m5-max-2026-05.jsonNext command
npm run bench:hygiene:doctor -- --clean-recording --output .local/benchmark-hygiene/current-reference-m5-max-2026-05.doctor.jsonMonitor command
npm run bench:hygiene:session -- monitor --clean-recording --label bitter-mill-current-reference-m5-max-2026-05 --sample-interval-ms 15000 --output data/ops/bitter-mill/incoming/current-reference-m5-max-2026-05/system-hygiene.jsonHygiene gate: Clean recording · clean intake root · memory metrics present · pre-existing swap warning-only · passing final snapshot · passing monitor samples · 0 swap I/O pages · no thermal/performance throttling
Gate: The hygiene doctor reports clean recording readiness before Bitter Mill is opened; if it reports excessive swap allocation, memory pressure, stale local inference processes, or thermal/performance warnings, stop before any inference run. / The all-runbooks readiness audit checks every Bitter Mill runbook with one captured machine and clean-recording system-hygiene snapshot before any runbook is opened.
Run M5 Max 128 GB frontier anchors under clean-recording lab hygiene
Use the owned M5 Max 128 GB plan to convert frontier and coding-agent watchlist signals into first-party Silicon Score lab rows, but only when the recording window proves zero swap I/O and clean system pressure.
8 models · 1 source
Status: Ready after hygiene · Run the listed clean-recording gate before benchmark execution, discovery, import, or publication changes.
Next command
npm run bench:lab:m5:anchors:dry-run -- --preflight --clean-recordingHygiene gate: Clean recording · memory metrics present · pre-existing swap warning-only · passing final snapshot · 0 swap I/O pages
Gate: Actual run must occur on a detected M5 Max 128 GB machine; runner mismatch warnings are blockers for canonical factory_measured promotion. / Output benchmark records resolve chip, RAM, model, runtime, quantization, source date, and generation tok/s before append review.
Monitor premier Apple Silicon benchmark sites for freshness gaps
Track LLMCheck, oMLX, asiai, Anubis OSS, apple-silicon-llm-bench, mac-llm-bench, and broad LocalScore discovery so Silicon Score notices frontier releases, benchmark hygiene patterns, and UX/evidence gaps before public rankings go stale.
7 sources
Next command
npm run research:compact -- read --url https://llmcheck.net/benchmarks --query "Qwen3.6 Gemma 4 Apple Silicon benchmark methodology raw measurements" --maxChars 2400 --refreshHygiene gate: Clean recording · pre-existing swap warning-only
Gate: Every candidate source fact is backed by a fetched URL or captured artifact. / Freshness-review reads use research:compact --refresh so currentness checks do not silently reuse stale cache entries.
Corroborate GLM-5.1 Apple Silicon viability
Official GLM-5.1 metadata and local-serving docs are captured, plus row-level oMLX and Hugging Face MLX quantization evidence on M3 Ultra 512GB; these are directional field signals until first-party reproduction captures clean-recording hygiene.
1 model · 7 signals · 5 sources
Next command
npm run research:compact -- read --url https://huggingface.co/zai-org/GLM-5.1 --query "local deployment quantized GGUF MLX KTransformers GLM-5.1 Apple Silicon" --maxChars 3000 --refreshHygiene gate: Clean recording · pre-existing swap warning-only
Gate: The official GLM-5.1 model metadata and release note are captured before public ranking copy treats it as current. / Any Apple Silicon row must explicitly identify GLM-5.1, hardware, runtime, quantization, context, and source date before it becomes a practitioner signal or benchmark candidate.
Corroborate Mistral Small 4 Apple Silicon runtime
Mistral Small 4 has official local-serving support, Hugging Face MLX conversion cards, LLMCheck trusted-reference Apple Silicon rows, and SharpAI HomeSec-Bench M5 Pro 64GB llama.cpp field reports. The HomeSec rows make it a concrete reproduction candidate, but they remain directional community/operator evidence until first-party Bitter Mill captures setup, quantization, context, methodology, and hygiene sidecars.
1 model · 3 signals · 6 sources
Next command
npm run research:compact -- read --url https://www.sharpai.org/benchmark/ --query "Mistral-Small-4-119B Q2_K_XL UD-IQ1_M MacBook Pro M5 Pro 64GB llama.cpp tok/s TTFT HomeSec-Bench" --maxChars 5000 --refreshHygiene gate: Clean recording · pre-existing swap warning-only
Gate: Any Apple Silicon row must explicitly identify Mistral Small 4, hardware, runtime, quantization, context or prompt shape, reported speed, and source date before it becomes a practitioner signal or benchmark candidate. / The SharpAI HomeSec-Bench M5 Pro 64GB rows are curated directional community_operator_report signals; use them to design first-party reproduction, not as canonical benchmark rows.
Corroborate Magistral Small Apple Silicon runtime
Magistral Small has official 24B reasoning-model grounding, a fit note that the quantized model can run within a 32GB RAM MacBook, and Hugging Face MLX conversion cards, but the 2026-05-05 compact refresh found no row-level Apple Silicon throughput evidence in LLMCheck, oMLX, mac-llm-bench, broad search, LocalLLaMA search, or the MLX cards. The existing M5 Max Bitter Mill batch owns first-party measurement; this task keeps external/runtime corroboration explicit without promoting fit guidance into speed evidence.
1 model · 1 signal · 5 sources
Next command
npm run research:compact -- read --url https://huggingface.co/mistralai/Magistral-Small-2506 --query "Magistral Small 2506 local inference quantized MacBook 32GB context degradation GGUF MLX Apple Silicon speed" --maxChars 4200 --refreshHygiene gate: Clean recording · pre-existing swap warning-only
Gate: Any Apple Silicon row must explicitly identify Magistral Small 2506, hardware, runtime, quantization, context or prompt shape, reported speed, and source date before it becomes a practitioner signal or benchmark candidate. / The official Hugging Face card and Mistral release post ground currentness, self-deployment, and runtime planning only; they do not establish Apple Silicon throughput.
Corroborate Llama 3.3 70B cross-tier coverage
Llama 3.3 70B now has Apple Silicon rows across M1, M2, M3, M4, and M5-era tiers, including trusted-reference LLMCheck anchors on M5 Max and M4 Ultra.
1 model · 4 sources
Next command
npm run research:compact -- read --url https://llmcheck.net/data/benchmarks.json --query "Llama 3.3 70B Apple Silicon chip ram engine tps date" --maxChars 2400 --refreshGate: Any candidate coverage row must explicitly identify Llama 3.3 70B, Apple Silicon hardware, RAM tier or clearly bounded chip tier, runtime, quantization, reported speed, and source date. / New community or competitor rows stay directional unless methodology and provenance are strong enough for trusted-reference treatment.
GLM-5
GLM-5. It is not yet published in the current frontier packs. Benchmark evidence includes 5 Apple Silicon benchmark rows. 1 official model brief captured. 7 fetched artifacts, 1 blocked, missing, or partial. 4 curated practitioner signals, 3 Apple Silicon-specific. GLM-5: 6 Apple Silicon field reports; best reported generation ~20 tok/s; best reported prompt processing ~187 tok/s; reported RAM use ~391.82-415.41GB; seen on M3 ULTRA, Mac Studio M3 ULTRA 512GB; via oMLX.
1 model · 2 signals · 1 source
Next command
npm run research:compact -- read --url https://z.ai/blog/glm-5 --query "GLM-5 launch local deployment vLLM SGLang Apple Silicon evidence" --maxChars 4000 --refreshGate: Keep the original practitioner claim low-confidence while the source capture_status=http_error. / The official Z.AI launch blog and BigModel docs may support GLM-5 currentness and local-serving planning, but they do not resolve the unavailable practitioner source and must not become Apple Silicon throughput evidence.
Reviewed: m4-max-40-core-gpu--qwen-3-32b--q4-k-m--lab
Keep the M4 Max lab value as first-party setup-specific evidence while treating Qwen 3 32B as a reference-only baseline displaced by newer Qwen3.5 and Qwen3.6 records.
Reviewed: m3-max-gpu-count-not-published--qwen-3-32b--q8-0--reddit-m3-max-64gb-api-sweep-20172-llamacpp
Keep the low long-context llama.cpp value as directional community evidence; it is useful for context-scaling caution but should not drive short-prompt recommendations.
Reviewed: m3-max-gpu-count-not-published--qwen-3-32b--q8-0--reddit-m3-max-64gb-api-sweep-20172-ollama
Keep the low long-context Ollama value as directional community evidence; it is useful for context-scaling caution but should not drive short-prompt recommendations.
Reviewed: m1-pro-16-core-gpu--llama-2-7b--q4-0--llamacpp
Source review did not recover a RAM tier for the M1 Pro 16-core GPU llama.cpp baseline row, so the benchmark remains useful but exact-machine confidence must stay limited.
Reviewed: m3-pro-18-core-gpu--llama-2-7b--q4-0--llamacpp
Source review did not recover a RAM tier for the M3 Pro 18-core GPU llama.cpp baseline row, so the benchmark remains useful but exact-machine confidence must stay limited.
Competitor currentness
These sources are tracked for frontier coverage, workflow UX, and benchmark hygiene. They inform research priorities; they do not bypass Silicon Score provenance or first-party reproduction.
Monitored surfaces
7
2 highest-priority Apple Silicon coverage sources
Weekly refresh loop
3
rerun with refreshed source reads before frontier conclusions move
Raw-data candidates
3
download or repository rows can enter import review after model grounding
Publication gates
7
competitor claims remain directional until the source-specific gate clears
highest
highest
high
high
medium
medium
medium
Evidence pressure
Metadata debt is now small. The bigger issue is still how much of the catalog depends on community rows and thin runtime coverage.
Macs
29
Models
55
Source captures
110
Unresolved issues
Reviewed: m4-max-40-core-gpu--qwen-3-32b--q4-k-m--lab
Keep the M4 Max lab value as first-party setup-specific evidence while treating Qwen 3 32B as a reference-only baseline displaced by newer Qwen3.5 and Qwen3.6 records.
Reviewed value kept
reviewed outlier kept
Reviewed: m3-max-gpu-count-not-published--qwen-3-32b--q8-0--reddit-m3-max-64gb-api-sweep-20172-llamacpp
Keep the low long-context llama.cpp value as directional community evidence; it is useful for context-scaling caution but should not drive short-prompt recommendations.
Reviewed value kept
reviewed outlier kept
Reviewed: m3-max-gpu-count-not-published--qwen-3-32b--q8-0--reddit-m3-max-64gb-api-sweep-20172-ollama
Keep the low long-context Ollama value as directional community evidence; it is useful for context-scaling caution but should not drive short-prompt recommendations.
Reviewed value kept
reviewed outlier kept
Reviewed: m1-pro-16-core-gpu--llama-2-7b--q4-0--llamacpp
Source review did not recover a RAM tier for the M1 Pro 16-core GPU llama.cpp baseline row, so the benchmark remains useful but exact-machine confidence must stay limited.
Reviewed source limit
source limit documented
Reviewed: m3-pro-18-core-gpu--llama-2-7b--q4-0--llamacpp
Source review did not recover a RAM tier for the M3 Pro 18-core GPU llama.cpp baseline row, so the benchmark remains useful but exact-machine confidence must stay limited.
Reviewed source limit
source limit documented
Reviewed: m4-max--qwen-3-32b--q4-k-m--estsauver-lm
Keep the LM Studio 10K-context value as trusted-reference evidence while preserving setup sensitivity in ranking explanations.
Reviewed value kept
reviewed outlier kept
Runtime taxonomy
Public runtime locks should stay simple. Backends and wrappers still matter, but they belong here as audit semantics rather than top-level navigation.
Llamafile
Llamafile wrapper on llama.cpp
llama.cpp stack
Public filter: llama.cpp stack.
MLX
MLX backend
MLX
Public filter: MLX.
Ollama
Ollama wrapper on llama.cpp
Ollama · llama.cpp stack
Public filter: Ollama · llama.cpp stack.
llama.cpp
llama.cpp backend
llama.cpp stack
Public filter: llama.cpp stack.
LM Studio
LM Studio wrapper on mixed
Audit only
Do not expose as a canonical runtime lock.
Methodology comparability
Frontier hotspots
Run Bitter Mill current-frontier batch on the owned M5 Max
22 frontier watchlist rows share this operator task: Qwen3.6-27B, Gemma 4 31B, Qwen3.5-27B, Qwen3.6-35B-A3B, Devstral Small 2 24B, Mistral Small 4 119B, and 16 more. Qwen3.6, MiniMax M2.7, Gemma 4 including E2B/E4B, Mistral Small 4, Ministral 3 compact models, Llama 4 Scout, gpt-oss 120B/20B, Nemotron Cascade 2, GLM-4.5-Air, Magistral Small, Devstral Small 2, and current Qwen coding/dense anchors should be reproduced from first-hand Bitter Mill runtime, quantization, and context sweeps before they become first-party measured evidence.
22 models · 1 source
Status: Ready after hygiene · Run the listed clean-recording gate before benchmark execution, discovery, import, or publication changes.
Next command
npm run bench:hygiene:doctor -- --clean-recording --output .local/benchmark-hygiene/current-frontier-m5-max-2026-05.doctor.jsonMonitor command
npm run bench:hygiene:session -- monitor --clean-recording --label bitter-mill-current-frontier-m5-max-2026-05 --sample-interval-ms 15000 --output data/ops/bitter-mill/incoming/current-frontier-m5-max-2026-05/system-hygiene.jsonHygiene gate: Clean recording · clean intake root · memory metrics present · pre-existing swap warning-only · passing final snapshot · passing monitor samples · 0 swap I/O pages · no thermal/performance throttling
Gate: The hygiene doctor reports clean recording readiness before Bitter Mill is opened; if it reports excessive swap allocation, memory pressure, stale local inference processes, or thermal/performance warnings, stop before any inference run. / The all-runbooks readiness audit checks every Bitter Mill runbook with one captured machine and clean-recording system-hygiene snapshot before any runbook is opened.
first party bitter mill batch
Qwen3.5-397B-A17B
Qwen3.5-397B-A17B. It appears across 4 lenses and 1 budget slice. Benchmark evidence includes 2 Apple Silicon benchmark rows. 1 official model brief captured. 8 fetched artifacts. 5 curated practitioner signals, 5 Apple Silicon-specific. Qwen3.5-397B-A17B: 4 Apple Silicon field reports; best reported generation ~30.81 tok/s; best reported prompt processing ~122.46 tok/s; seen on MacBook Pro M5 MAX 128GB; via llama.cpp, flash-moe, MLX.
1 model · 1 source
Status: Ready after hygiene · Run the listed clean-recording gate before benchmark execution, discovery, import, or publication changes.
Next command
npm run bench:hygiene:doctor -- --clean-recording --output .local/benchmark-hygiene/import-bitter-mill-qwen-3-5-397b-a17b.doctor.jsonMonitor command
npm run bench:hygiene:session -- monitor --clean-recording --label import-bitter-mill-qwen-3-5-397b-a17b --sample-interval-ms 15000 --output data/ops/bitter-mill/incoming/qwen-3-5-397b-a17b.system-hygiene.jsonHygiene gate: Clean recording · memory metrics present · pre-existing swap warning-only · passing final snapshot · passing monitor samples · 0 swap I/O pages · no thermal/performance throttling
Gate: The hygiene doctor reports clean recording readiness before Bitter Mill is opened; if it reports excessive swap allocation, memory pressure, stale local inference processes, or thermal/performance warnings, stop before any inference run. / The hygiene monitor starts before Bitter Mill is opened and stops only after the model export lands; a same-base monitor hygiene sidecar exists beside the Bitter Mill export.
first party bitter mill import
Qwen3.5-122B-A10B
Qwen3.5-122B-A10B. It appears across 4 lenses and 3 budget slices. Benchmark evidence includes 6 Apple Silicon benchmark rows. 1 official model brief captured. 15 fetched artifacts, 1 blocked, missing, or partial. 16 curated practitioner signals, 16 Apple Silicon-specific. Qwen3.5-122B-A10B: 24 Apple Silicon field reports; best reported generation ~65.853 tok/s; best reported prompt processing ~1239.734 tok/s; reported RAM use ~71.91-102GB; seen on MacBook Pro M5 MAX 128GB, Mac Studio M3 ULTRA 256GB, Mac Studio M4 MAX 128GB; via MLX, oMLX, EXO over Thunderbolt 5 RDMA, llama.cpp.
1 model · 1 source
Status: Ready after hygiene · Run the listed clean-recording gate before benchmark execution, discovery, import, or publication changes.
Next command
npm run bench:hygiene:doctor -- --clean-recording --output .local/benchmark-hygiene/import-bitter-mill-qwen-3-5-122b-a10b.doctor.jsonMonitor command
npm run bench:hygiene:session -- monitor --clean-recording --label import-bitter-mill-qwen-3-5-122b-a10b --sample-interval-ms 15000 --output data/ops/bitter-mill/incoming/qwen-3-5-122b-a10b.system-hygiene.jsonHygiene gate: Clean recording · memory metrics present · pre-existing swap warning-only · passing final snapshot · passing monitor samples · 0 swap I/O pages · no thermal/performance throttling
Gate: The hygiene doctor reports clean recording readiness before Bitter Mill is opened; if it reports excessive swap allocation, memory pressure, stale local inference processes, or thermal/performance warnings, stop before any inference run. / The hygiene monitor starts before Bitter Mill is opened and stops only after the model export lands; a same-base monitor hygiene sidecar exists beside the Bitter Mill export.
first party bitter mill import
GLM-5.1
GLM-5.1. It is not yet published in the current frontier packs. No canonical Apple Silicon benchmark rows yet; field speed reports are still directional. 3 official model briefs captured. 13 fetched artifacts. 7 curated practitioner signals, 6 Apple Silicon-specific. GLM-5.1: 7 Apple Silicon field reports; best reported generation ~19.527 tok/s; best reported prompt processing ~194.216 tok/s; reported RAM use ~251-382.6GB; seen on Mac Studio M3 ULTRA 512GB, Mac Studio M3 ULTRA 256GB; via MLX, oMLX.
1 model · 7 signals · 5 sources
Next command
npm run research:compact -- read --url https://huggingface.co/zai-org/GLM-5.1 --query "local deployment quantized GGUF MLX KTransformers GLM-5.1 Apple Silicon" --maxChars 3000 --refreshHygiene gate: Clean recording · pre-existing swap warning-only
Gate: The official GLM-5.1 model metadata and release note are captured before public ranking copy treats it as current. / Any Apple Silicon row must explicitly identify GLM-5.1, hardware, runtime, quantization, context, and source date before it becomes a practitioner signal or benchmark candidate.
coverage expansion
Run M5 Max 128 GB frontier anchors under clean-recording lab hygiene
4 frontier watchlist rows share this operator task: Qwen3-Coder-30B-A3B, Llama 3.3 70B, Qwen 3 30B-A3B, Gemma 3 27B. Use the owned M5 Max 128 GB plan to convert frontier and coding-agent watchlist signals into first-party Silicon Score lab rows, but only when the recording window proves zero swap I/O and clean system pressure.
8 models · 1 source
Status: Ready after hygiene · Run the listed clean-recording gate before benchmark execution, discovery, import, or publication changes.
Next command
npm run bench:lab:m5:anchors:dry-run -- --preflight --clean-recordingHygiene gate: Clean recording · memory metrics present · pre-existing swap warning-only · passing final snapshot · 0 swap I/O pages
Gate: Actual run must occur on a detected M5 Max 128 GB machine; runner mismatch warnings are blockers for canonical factory_measured promotion. / Output benchmark records resolve chip, RAM, model, runtime, quantization, source date, and generation tok/s before append review.
lab verification
GLM-5
GLM-5. It is not yet published in the current frontier packs. Benchmark evidence includes 5 Apple Silicon benchmark rows. 1 official model brief captured. 7 fetched artifacts, 1 blocked, missing, or partial. 4 curated practitioner signals, 3 Apple Silicon-specific. GLM-5: 6 Apple Silicon field reports; best reported generation ~20 tok/s; best reported prompt processing ~187 tok/s; reported RAM use ~391.82-415.41GB; seen on M3 ULTRA, Mac Studio M3 ULTRA 512GB; via oMLX.
1 model · 2 signals · 1 source
Next command
npm run research:compact -- read --url https://z.ai/blog/glm-5 --query "GLM-5 launch local deployment vLLM SGLang Apple Silicon evidence" --maxChars 4000 --refreshGate: Keep the original practitioner claim low-confidence while the source capture_status=http_error. / The official Z.AI launch blog and BigModel docs may support GLM-5 currentness and local-serving planning, but they do not resolve the unavailable practitioner source and must not become Apple Silicon throughput evidence.
practitioner source recovery
Run Bitter Mill current-reference batch on the owned M5 Max
11 frontier watchlist rows share this operator task: Devstral Small 1.1, Mistral Small 3.1 24B, Qwen 3 235B-A22B, DeepSeek R1 Distill Qwen 32B, GLM-4.7-Flash, DeepSeek R1 Distill Llama 70B, and 5 more. Qwen 3 32B, Qwen 3 235B-A22B, Devstral Small 1.1, Qwen 3 8B, Nemotron-3-Nano, GLM-4.7-Flash, Phi-4 14B, Mistral Small 3.1, DeepSeek R1 distills, and Qwen 2.5 72B have Apple Silicon evidence or practitioner signals but should be treated as reference/baseline reproduction work rather than the current frontier lane. Qwen3.5-35B-A3B stays in the current-frontier M5 batch as the practical Qwen3.6 displacement comparator.
12 models · 1 source
Status: Ready after hygiene · Run the listed clean-recording gate before benchmark execution, discovery, import, or publication changes.
Next command
npm run bench:hygiene:doctor -- --clean-recording --output .local/benchmark-hygiene/current-reference-m5-max-2026-05.doctor.jsonMonitor command
npm run bench:hygiene:session -- monitor --clean-recording --label bitter-mill-current-reference-m5-max-2026-05 --sample-interval-ms 15000 --output data/ops/bitter-mill/incoming/current-reference-m5-max-2026-05/system-hygiene.jsonHygiene gate: Clean recording · clean intake root · memory metrics present · pre-existing swap warning-only · passing final snapshot · passing monitor samples · 0 swap I/O pages · no thermal/performance throttling
Gate: The hygiene doctor reports clean recording readiness before Bitter Mill is opened; if it reports excessive swap allocation, memory pressure, stale local inference processes, or thermal/performance warnings, stop before any inference run. / The all-runbooks readiness audit checks every Bitter Mill runbook with one captured machine and clean-recording system-hygiene snapshot before any runbook is opened.
first party bitter mill batch
Operator queue
Run Bitter Mill current-frontier batch on the owned M5 Max
Qwen3.6, MiniMax M2.7, Gemma 4 including E2B/E4B, Mistral Small 4, Ministral 3 compact models, Llama 4 Scout, gpt-oss 120B/20B, Nemotron Cascade 2, GLM-4.5-Air, Magistral Small, Devstral Small 2, and current Qwen coding/dense anchors should be reproduced from first-hand Bitter Mill runtime, quantization, and context sweeps before they become first-party measured evidence.
22 models · 1 source
Status: Ready after hygiene · Run the listed clean-recording gate before benchmark execution, discovery, import, or publication changes.
Runbook
data/ops/bitter-mill/current-frontier-m5-max-2026-05.jsonNext command
npm run bench:hygiene:doctor -- --clean-recording --output .local/benchmark-hygiene/current-frontier-m5-max-2026-05.doctor.jsonMonitor command
npm run bench:hygiene:session -- monitor --clean-recording --label bitter-mill-current-frontier-m5-max-2026-05 --sample-interval-ms 15000 --output data/ops/bitter-mill/incoming/current-frontier-m5-max-2026-05/system-hygiene.jsonHygiene gate: Clean recording · clean intake root · memory metrics present · pre-existing swap warning-only · passing final snapshot · passing monitor samples · 0 swap I/O pages · no thermal/performance throttling
Gate: The hygiene doctor reports clean recording readiness before Bitter Mill is opened; if it reports excessive swap allocation, memory pressure, stale local inference processes, or thermal/performance warnings, stop before any inference run. / The all-runbooks readiness audit checks every Bitter Mill runbook with one captured machine and clean-recording system-hygiene snapshot before any runbook is opened.
first party bitter mill batch
Run planned M4 Ultra 256GB frontier arrival batch
When the 256GB Mac Studio arrives, run a clean-recording Bitter Mill sweep across the active M5 current-frontier target set including MiniMax M2.7 plus high-memory and fit-boundary extras, runtimes, quantizations including current MLX dynamic low-bit profiles, and context ladders before changing Silicon Score's featured environment.
26 models · 1 source
Status: Blocked until arrival · M4 Ultra 256 GB stays queued until the machine is physically present, hardware identity is verified, and clean-recording preflight passes.
Runbook
data/ops/bitter-mill/planned/m4-ultra-256gb-frontier-2026-06.jsonNext command
npm run bench:hygiene:doctor -- --clean-recording --output .local/benchmark-hygiene/m4-ultra-256gb-frontier-2026-06.doctor.jsonMonitor command
npm run bench:hygiene:session -- monitor --clean-recording --label bitter-mill-m4-ultra-256gb-frontier-2026-06 --sample-interval-ms 15000 --output data/ops/bitter-mill/incoming/m4-ultra-256gb-frontier-2026-06/system-hygiene.jsonHygiene gate: Clean recording · clean intake root · memory metrics present · pre-existing swap warning-only · passing final snapshot · passing monitor samples · 0 swap I/O pages · no thermal/performance throttling
Gate: The hygiene doctor reports clean recording readiness before Bitter Mill is opened on the arrived Mac Studio. / The active all-runbooks readiness audit passes for checked-in active runbooks, then the planned M4 Ultra runbook's explicit check passes on the arrived Mac Studio.
first party bitter mill batch
Prepare M4 Ultra 256GB Mac Studio as the next featured lab environment
In June 2026 the local lab is expected to receive a 256GB Mac Studio provisionally labeled M4 Ultra, replacing the M5 Max 128GB MacBook Pro as the featured environment only after arrival, hardware detection, and clean first-party evidence justify changing the public default.
1 source
Status: Blocked until arrival · M4 Ultra 256 GB stays queued until the machine is physically present, hardware identity is verified, and clean-recording preflight passes.
Next command
npm run validate:lab-environmentsHygiene gate: Clean recording · pre-existing swap warning-only · passing final snapshot · 0 swap I/O pages · no thermal/performance throttling
Gate: data/ops/lab-environments.json keeps the M5 Max 128GB MacBook Pro as current active featured environment until the M4 Ultra 256GB Mac Studio is physically present. / Arrival capture records system_profiler chip name, GPU cores, model identifier, and 256GB unified memory before any featured_local_benchmark or FEATURED_MACHINE_ID change.
lab environment transition
Upgrade Qwen3.5-122B-A10B reference rows with Bitter Mill runs
Qwen3.5-122B-A10B is a high-memory frontier MoE candidate with trusted-reference rows and field speed reports, and this queued Bitter Mill import will test whether it belongs in high-end Mac advice.
1 model · 1 source
Status: Ready after hygiene · Run the listed clean-recording gate before benchmark execution, discovery, import, or publication changes.
Next command
npm run bench:hygiene:doctor -- --clean-recording --output .local/benchmark-hygiene/import-bitter-mill-qwen-3-5-122b-a10b.doctor.jsonMonitor command
npm run bench:hygiene:session -- monitor --clean-recording --label import-bitter-mill-qwen-3-5-122b-a10b --sample-interval-ms 15000 --output data/ops/bitter-mill/incoming/qwen-3-5-122b-a10b.system-hygiene.jsonHygiene gate: Clean recording · memory metrics present · pre-existing swap warning-only · passing final snapshot · passing monitor samples · 0 swap I/O pages · no thermal/performance throttling
Gate: The hygiene doctor reports clean recording readiness before Bitter Mill is opened; if it reports excessive swap allocation, memory pressure, stale local inference processes, or thermal/performance warnings, stop before any inference run. / The hygiene monitor starts before Bitter Mill is opened and stops only after the model export lands; a same-base monitor hygiene sidecar exists beside the Bitter Mill export.
first party bitter mill import
Upgrade Qwen3.5-397B-A17B reference rows with Bitter Mill runs
Qwen3.5-397B-A17B is the current high-memory Qwen3.5 flagship with sparse Apple Silicon reference evidence, and this queued Bitter Mill import keeps it behind clean first-party reproduction.
1 model · 1 source
Status: Ready after hygiene · Run the listed clean-recording gate before benchmark execution, discovery, import, or publication changes.
Next command
npm run bench:hygiene:doctor -- --clean-recording --output .local/benchmark-hygiene/import-bitter-mill-qwen-3-5-397b-a17b.doctor.jsonMonitor command
npm run bench:hygiene:session -- monitor --clean-recording --label import-bitter-mill-qwen-3-5-397b-a17b --sample-interval-ms 15000 --output data/ops/bitter-mill/incoming/qwen-3-5-397b-a17b.system-hygiene.jsonHygiene gate: Clean recording · memory metrics present · pre-existing swap warning-only · passing final snapshot · passing monitor samples · 0 swap I/O pages · no thermal/performance throttling
Gate: The hygiene doctor reports clean recording readiness before Bitter Mill is opened; if it reports excessive swap allocation, memory pressure, stale local inference processes, or thermal/performance warnings, stop before any inference run. / The hygiene monitor starts before Bitter Mill is opened and stops only after the model export lands; a same-base monitor hygiene sidecar exists beside the Bitter Mill export.
first party bitter mill import
Run Bitter Mill current-reference batch on the owned M5 Max
Qwen 3 32B, Qwen 3 235B-A22B, Devstral Small 1.1, Qwen 3 8B, Nemotron-3-Nano, GLM-4.7-Flash, Phi-4 14B, Mistral Small 3.1, DeepSeek R1 distills, and Qwen 2.5 72B have Apple Silicon evidence or practitioner signals but should be treated as reference/baseline reproduction work rather than the current frontier lane. Qwen3.5-35B-A3B stays in the current-frontier M5 batch as the practical Qwen3.6 displacement comparator.
12 models · 1 source
Status: Ready after hygiene · Run the listed clean-recording gate before benchmark execution, discovery, import, or publication changes.
Runbook
data/ops/bitter-mill/current-reference-m5-max-2026-05.jsonNext command
npm run bench:hygiene:doctor -- --clean-recording --output .local/benchmark-hygiene/current-reference-m5-max-2026-05.doctor.jsonMonitor command
npm run bench:hygiene:session -- monitor --clean-recording --label bitter-mill-current-reference-m5-max-2026-05 --sample-interval-ms 15000 --output data/ops/bitter-mill/incoming/current-reference-m5-max-2026-05/system-hygiene.jsonHygiene gate: Clean recording · clean intake root · memory metrics present · pre-existing swap warning-only · passing final snapshot · passing monitor samples · 0 swap I/O pages · no thermal/performance throttling
Gate: The hygiene doctor reports clean recording readiness before Bitter Mill is opened; if it reports excessive swap allocation, memory pressure, stale local inference processes, or thermal/performance warnings, stop before any inference run. / The all-runbooks readiness audit checks every Bitter Mill runbook with one captured machine and clean-recording system-hygiene snapshot before any runbook is opened.
first party bitter mill batch
Run M5 Max 128 GB frontier anchors under clean-recording lab hygiene
Use the owned M5 Max 128 GB plan to convert frontier and coding-agent watchlist signals into first-party Silicon Score lab rows, but only when the recording window proves zero swap I/O and clean system pressure.
8 models · 1 source
Status: Ready after hygiene · Run the listed clean-recording gate before benchmark execution, discovery, import, or publication changes.
Next command
npm run bench:lab:m5:anchors:dry-run -- --preflight --clean-recordingHygiene gate: Clean recording · memory metrics present · pre-existing swap warning-only · passing final snapshot · 0 swap I/O pages
Gate: Actual run must occur on a detected M5 Max 128 GB machine; runner mismatch warnings are blockers for canonical factory_measured promotion. / Output benchmark records resolve chip, RAM, model, runtime, quantization, source date, and generation tok/s before append review.
lab verification
Monitor premier Apple Silicon benchmark sites for freshness gaps
Track LLMCheck, oMLX, asiai, Anubis OSS, apple-silicon-llm-bench, mac-llm-bench, and broad LocalScore discovery so Silicon Score notices frontier releases, benchmark hygiene patterns, and UX/evidence gaps before public rankings go stale.
7 sources
Next command
npm run research:compact -- read --url https://llmcheck.net/benchmarks --query "Qwen3.6 Gemma 4 Apple Silicon benchmark methodology raw measurements" --maxChars 2400 --refreshHygiene gate: Clean recording · pre-existing swap warning-only
Gate: Every candidate source fact is backed by a fetched URL or captured artifact. / Freshness-review reads use research:compact --refresh so currentness checks do not silently reuse stale cache entries.
competitor freshness review
Corroborate GLM-5.1 Apple Silicon viability
Official GLM-5.1 metadata and local-serving docs are captured, plus row-level oMLX and Hugging Face MLX quantization evidence on M3 Ultra 512GB; these are directional field signals until first-party reproduction captures clean-recording hygiene.
1 model · 7 signals · 5 sources
Next command
npm run research:compact -- read --url https://huggingface.co/zai-org/GLM-5.1 --query "local deployment quantized GGUF MLX KTransformers GLM-5.1 Apple Silicon" --maxChars 3000 --refreshHygiene gate: Clean recording · pre-existing swap warning-only
Gate: The official GLM-5.1 model metadata and release note are captured before public ranking copy treats it as current. / Any Apple Silicon row must explicitly identify GLM-5.1, hardware, runtime, quantization, context, and source date before it becomes a practitioner signal or benchmark candidate.
coverage expansion
Corroborate Mistral Small 4 Apple Silicon runtime
Mistral Small 4 has official local-serving support, Hugging Face MLX conversion cards, LLMCheck trusted-reference Apple Silicon rows, and SharpAI HomeSec-Bench M5 Pro 64GB llama.cpp field reports. The HomeSec rows make it a concrete reproduction candidate, but they remain directional community/operator evidence until first-party Bitter Mill captures setup, quantization, context, methodology, and hygiene sidecars.
1 model · 3 signals · 6 sources
Next command
npm run research:compact -- read --url https://www.sharpai.org/benchmark/ --query "Mistral-Small-4-119B Q2_K_XL UD-IQ1_M MacBook Pro M5 Pro 64GB llama.cpp tok/s TTFT HomeSec-Bench" --maxChars 5000 --refreshHygiene gate: Clean recording · pre-existing swap warning-only
Gate: Any Apple Silicon row must explicitly identify Mistral Small 4, hardware, runtime, quantization, context or prompt shape, reported speed, and source date before it becomes a practitioner signal or benchmark candidate. / The SharpAI HomeSec-Bench M5 Pro 64GB rows are curated directional community_operator_report signals; use them to design first-party reproduction, not as canonical benchmark rows.
coverage expansion
Corroborate Magistral Small Apple Silicon runtime
Magistral Small has official 24B reasoning-model grounding, a fit note that the quantized model can run within a 32GB RAM MacBook, and Hugging Face MLX conversion cards, but the 2026-05-05 compact refresh found no row-level Apple Silicon throughput evidence in LLMCheck, oMLX, mac-llm-bench, broad search, LocalLLaMA search, or the MLX cards. The existing M5 Max Bitter Mill batch owns first-party measurement; this task keeps external/runtime corroboration explicit without promoting fit guidance into speed evidence.
1 model · 1 signal · 5 sources
Next command
npm run research:compact -- read --url https://huggingface.co/mistralai/Magistral-Small-2506 --query "Magistral Small 2506 local inference quantized MacBook 32GB context degradation GGUF MLX Apple Silicon speed" --maxChars 4200 --refreshHygiene gate: Clean recording · pre-existing swap warning-only
Gate: Any Apple Silicon row must explicitly identify Magistral Small 2506, hardware, runtime, quantization, context or prompt shape, reported speed, and source date before it becomes a practitioner signal or benchmark candidate. / The official Hugging Face card and Mistral release post ground currentness, self-deployment, and runtime planning only; they do not establish Apple Silicon throughput.
coverage expansion
Corroborate Llama 3.3 70B cross-tier coverage
Llama 3.3 70B now has Apple Silicon rows across M1, M2, M3, M4, and M5-era tiers, including trusted-reference LLMCheck anchors on M5 Max and M4 Ultra.
1 model · 4 sources
Next command
npm run research:compact -- read --url https://llmcheck.net/data/benchmarks.json --query "Llama 3.3 70B Apple Silicon chip ram engine tps date" --maxChars 2400 --refreshGate: Any candidate coverage row must explicitly identify Llama 3.3 70B, Apple Silicon hardware, RAM tier or clearly bounded chip tier, runtime, quantization, reported speed, and source date. / New community or competitor rows stay directional unless methodology and provenance are strong enough for trusted-reference treatment.
coverage expansion
Model scorecards
These standardized Apple Silicon evals add task-shape truth to the catalog. They come from one fixed high-end Mac and runtime path, so they should not override the main rankings, but they are excellent for understanding which models are fast, balanced, coding-heavy, or tool-soft.
Qwen3.5-122B-A10B
vLLM-MLX
Highest overall quality in this standardized set, but it demands real memory.
vLLM-MLX SCORECARD.md · discussion · 2026-03-04
Qwen3.5-122B-A10B
vLLM-MLX
The best value version in this scorecard: near-frontier quality at roughly half the RAM.
vLLM-MLX SCORECARD.md · discussion · 2026-03-04
Qwen3.5-35B-A3B
vLLM-MLX
The stronger version of the 35B MoE story: fast and much more balanced.
vLLM-MLX SCORECARD.md · discussion · 2026-03-04
Qwen3-Coder-Next
vLLM-MLX
Slightly slower than 4-bit, but reasoning is stronger and coding stays high.
vLLM-MLX SCORECARD.md · discussion · 2026-03-04
Qwen3-Coder-Next
vLLM-MLX
The fast coding-first option in this scorecard, with strong tool behavior.
vLLM-MLX SCORECARD.md · discussion · 2026-03-04
GLM-4.5-Air
vLLM-MLX
More balanced than the flash variant, but materially heavier.
vLLM-MLX SCORECARD.md · discussion · 2026-03-04
Qwen3.5-27B
vLLM-MLX
A strong fits-anywhere coding and tool-use compromise.
vLLM-MLX SCORECARD.md · discussion · 2026-03-04
Qwen3.5-35B-A3B
vLLM-MLX
Very fast for its size, but reasoning softness is visible in the standardized tasks.
vLLM-MLX SCORECARD.md · discussion · 2026-03-04
Qwen3.5-9B
vLLM-MLX
The smallest model in this set that still looks broadly useful for agent-style work.
vLLM-MLX SCORECARD.md · discussion · 2026-03-04
Devstral Small 2 24B
vLLM-MLX
Strong coding score, but tool calling is poor in this standardized setup.
vLLM-MLX SCORECARD.md · discussion · 2026-03-04
Workflow metrics
These records capture effective throughput and prefill-heavy scenarios from the same Mac and model across different runtime paths. They are valuable for teaching runtime choice, but they should stay in the audit lane rather than flatten into headline tokens-per-second rows.
Qwen 3 30B-A3B
Effective tok/s · Interactive
2026-03-20
These are effective throughput figures on a multi-turn ops-agent scenario. They include prefill and wrapper behavior, so they should teach runtime choice, not replace decode-speed benchmark rows.
LM Studio 41.7 · llama.cpp 41.4 · oMLX 38.0 · Ollama 26.0
Qwen3.5-35B-A3B
Effective tok/s · Interactive
2026-03-20
These are effective throughput results on an ops-agent workflow. They are best used to compare runtime behavior and caching quality on Apple Silicon, not to replace canonical decode-speed rows.
oMLX 38.0 · Rapid-MLX 35.6 · mlx-openai-server 26.2 · LM Studio (GGUF) 17.6 · LM Studio (MLX) 17.0
Qwen3.5-35B-A3B
Effective tok/s · 8,000 ctx
2026-03-20
This is an 8K prefill-stress comparison. It is useful for understanding caching and long-context behavior, not for headline decode-speed ranking.
oMLX 16.4 · mlx-openai-server 8.7 · Rapid-MLX 8.5 · LM Studio (GGUF) 7.8 · LM Studio (MLX) 5.9
Qwen 3 30B-A3B
Effective tok/s · 8,000 ctx
2026-03-20
This compares wrappers and backends on an 8K prefill-stress scenario. It is useful for long-context teaching, but it is not a canonical decode-speed row.
MLX fp16 8.6 · GGUF 7.6 · MLX bf16 6.0
Research map
Not every source should influence the product equally. This map makes those roles explicit so rankings can say when an answer is measured, estimated, or still fit-first.
First-party
Where measured truth should be upgraded into canon.
Bitter Mill local inference traces
first party local inference engine
Silicon Score Lab
first party lab
Benchmark reference
Where comparable benchmark anchors and runtime-change signals usually appear first.
Anubis OSS Apple Silicon benchmarks
competitor benchmark tool
apple-silicon-llm-bench
independent benchmark publication
asiai Apple Silicon benchmarks
competitor benchmark tool
Awni Hannun benchmark gists
maintainer benchmark gists
famstack.dev benchmark guides
independent benchmark publication
Hugging Face model hub
model registry
LLMCheck Apple Silicon benchmarks
competitor benchmark site
mac-llm-bench
independent benchmark publication
mlx-lm pull requests
maintainer pull requests
oMLX community benchmarks
competitor benchmark site
oMLX repository
runtime repo
SharpAI HomeSec-Bench
published benchmark page
vLLM-MLX repository
runtime repo
Zach Rattner benchmark gists
community benchmark gists
Practitioner
Where operators reveal workflow reality and caveats before the benchmark layer catches up.
llama.cpp GitHub discussions
maintainer discussion
Reddit /r/LocalLLaMA
operator forum
Discovery
Where release movement starts, but not where performance truth should harden.
LocalScore accelerator runs
community benchmark aggregator
X local-AI chatter
social discovery