← Canonical rankings

llama-3-2-1b-instruct Apple Silicon benchmark and best Macs

Start with the ranked Mac table above, then audit tokens per second, RAM fit, quantization, runtimes, and source links before trusting the model on your Mac.

Quantizations observed: Q4_K - Medium

63Benchmark rows
63Chip tiers covered
229.0Fastest avg tok/s (M5 Max (32-core GPU, 36 GB))
Minimum RAM observed

Fastest published result is 229.0 tok/s on M5 Max (32-core GPU, 36 GB) at Q4_K - Medium. Published runtimes include llamafile. Start with Rankings for the decision, then use the raw rows below to audit the evidence.

Based on 63 external benchmarks; no lab runs yet.

Published runtimes: llamafile.

How fast is llama-3-2-1b-instruct on Mac?

llama-3-2-1b-instruct currently has a fastest published Mac result of 229.0 tok/s on M5 Max (32-core GPU, 36 GB) at Q4_K - Medium. Check the raw rows on this page before comparing that number to a different runtime, quantization, or context length.

What is the best Mac for llama-3-2-1b-instruct?

There is no direct best-Mac benchmark winner for llama-3-2-1b-instruct yet. Use the ranking above to find the current Apple Silicon fit path and return to this page when direct rows land.

Does llama-3-2-1b-instruct fit on Apple Silicon?

llama-3-2-1b-instruct does not yet have a published minimum-RAM fit row on this page. Use Fit before assuming it will run cleanly on a specific Mac.

Published chip coverage includes M5 Max (32-core GPU, 36 GB), M4 Max (40-core GPU, 128 GB), M4 Max (40-core GPU, 64 GB), M4 Max (40-core GPU, 48 GB), M3 Ultra (80-core GPU, 512 GB) plus 58 more chip tiers. Fastest published row is 229.0 tok/s on M5 Max (32-core GPU, 36 GB) at Q4_K - Medium.

Raw benchmark rows for llama-3-2-1b-instruct

Rows stay below the ranking because this page is answer-first. Use them to inspect exact chips, quantizations, runtimes, and sources.

ChipQuantRAM req.ContextAvg tok/sPrompt tok/sRuntimeSource
M5 Max (32-core GPU, 36 GB)Q4_K - Medium229.0 tok/s3620.7 tok/sllamafileref
M4 Max (40-core GPU, 128 GB)Q4_K - Medium182.6 tok/s3833.8 tok/sllamafileref
M4 Max (40-core GPU, 64 GB)Q4_K - Medium180.2 tok/s3857.5 tok/sllamafileref
M4 Max (40-core GPU, 48 GB)Q4_K - Medium179.0 tok/s3750.6 tok/sllamafileref
M3 Ultra (80-core GPU, 512 GB)Q4_K - Medium178.8 tok/s5601.0 tok/sllamafileref
M3 Ultra (80-core GPU, 256 GB)Q4_K - Medium177.9 tok/s4999.5 tok/sllamafileref
M2 Ultra (60-core GPU, 128 GB)Q4_K - Medium176.4 tok/s3295.5 tok/sllamafileref
M2 Ultra (60-core GPU, 64 GB)Q4_K - Medium174.1 tok/s3290.2 tok/sllamafileref
M2 Ultra (60-core GPU, 192 GB)Q4_K - Medium169.8 tok/s3272.2 tok/sllamafileref
M4 Max (32-core GPU, 36 GB)Q4_K - Medium166.5 tok/s3268.9 tok/sllamafileref
M4 Max (GPU count not published, 128 GB)Q4_K - Medium156.3 tok/s679.7 tok/sllamafileref
M2 Max (38-core GPU, 32 GB)Q4_K - Medium153.0 tok/s2551.4 tok/sllamafileref
M1 Ultra (64-core GPU, 128 GB)Q4_K - Medium151.1 tok/s2977.1 tok/sllamafileref
M3 Max (40-core GPU, 48 GB)Q4_K - Medium149.0 tok/s3399.1 tok/sllamafileref
M3 Max (40-core GPU, 128 GB)Q4_K - Medium146.3 tok/s3291.9 tok/sllamafileref
M1 Ultra (48-core GPU, 128 GB)Q4_K - Medium138.0 tok/s2582.4 tok/sllamafileref
M3 Max (30-core GPU, 36 GB)Q4_K - Medium133.0 tok/s2553.3 tok/sllamafileref
M3 Max (30-core GPU, 96 GB)Q4_K - Medium132.9 tok/s2520.0 tok/sllamafileref
M2 Max (30-core GPU, 32 GB)Q4_K - Medium127.6 tok/s2031.7 tok/sllamafileref
M1 Max (32-core GPU, 32 GB)Q4_K - Medium125.8 tok/s2025.2 tok/sllamafileref
M1 Max (32-core GPU, 64 GB)Q4_K - Medium120.7 tok/s1972.1 tok/sllamafileref
M2 Ultra (GPU count not published, 128 GB)Q4_K - Medium120.4 tok/s659.8 tok/sllamafileref
M4 Pro (20-core GPU, 24 GB)Q4_K - Medium119.2 tok/s2128.8 tok/sllamafileref
M4 Pro (20-core GPU, 48 GB)Q4_K - Medium118.9 tok/s2134.1 tok/sllamafileref
M4 Pro (20-core GPU, 64 GB)Q4_K - Medium118.6 tok/s2145.1 tok/sllamafileref
M4 Pro (16-core GPU, 64 GB)Q4_K - Medium111.9 tok/s1858.9 tok/sllamafileref
M4 Pro (16-core GPU, 48 GB)Q4_K - Medium111.0 tok/s1754.6 tok/sllamafileref
M4 Pro (16-core GPU, 24 GB)Q4_K - Medium110.9 tok/s1823.8 tok/sllamafileref
M3 Max (40-core GPU, 64 GB)Q4_K - Medium107.0 tok/s2521.6 tok/sllamafileref
M1 Max (24-core GPU, 64 GB)Q4_K - Medium105.7 tok/s1669.5 tok/sllamafileref
M2 Pro (19-core GPU, 32 GB)Q4_K - Medium100.3 tok/s1487.6 tok/sllamafileref
M2 Pro (19-core GPU, 16 GB)Q4_K - Medium99.5 tok/s1457.9 tok/sllamafileref
M5 (10-core GPU, 32 GB)Q4_K - Medium98.4 tok/s1271.5 tok/sllamafileref
M5 (10-core GPU, 16 GB)Q4_K - Medium98.1 tok/s1244.9 tok/sllamafileref
M1 Max (24-core GPU, 32 GB)Q4_K - Medium93.9 tok/s1582.0 tok/sllamafileref
M2 Pro (16-core GPU, 32 GB)Q4_K - Medium91.5 tok/s1281.4 tok/sllamafileref
M2 Pro (16-core GPU, 16 GB)Q4_K - Medium91.1 tok/s1328.2 tok/sllamafileref
M3 Pro (18-core GPU, 36 GB)Q4_K - Medium89.8 tok/s1586.2 tok/sllamafileref
M3 Pro (14-core GPU, 36 GB)Q4_K - Medium88.2 tok/s1327.9 tok/sllamafileref
M3 Pro (14-core GPU, 18 GB)Q4_K - Medium88.1 tok/s1344.0 tok/sllamafileref
M3 Pro (18-core GPU, 18 GB)Q4_K - Medium85.6 tok/s1573.7 tok/sllamafileref
M1 Pro (16-core GPU, 16 GB)Q4_K - Medium78.2 tok/s1158.1 tok/sllamafileref
M1 Pro (16-core GPU, 32 GB)Q4_K - Medium77.2 tok/s1166.0 tok/sllamafileref
M4 (10-core GPU, 16 GB)Q4_K - Medium76.2 tok/s1091.1 tok/sllamafileref
M4 (10-core GPU, 32 GB)Q4_K - Medium75.6 tok/s1069.9 tok/sllamafileref
M4 (10-core GPU, 24 GB)Q4_K - Medium75.4 tok/s1036.0 tok/sllamafileref
M1 Pro (14-core GPU, 16 GB)Q4_K - Medium71.8 tok/s1040.1 tok/sllamafileref
M1 Pro (14-core GPU, 32 GB)Q4_K - Medium71.0 tok/s1063.8 tok/sllamafileref
M4 (GPU count not published, 16 GB)Q4_K - Medium68.0 tok/s239.0 tok/sllamafileref
M3 (10-core GPU, 16 GB)Q4_K - Medium67.2 tok/s931.5 tok/sllamafileref
M4 (8-core GPU, 16 GB)Q4_K - Medium65.9 tok/s897.0 tok/sllamafileref
M3 (10-core GPU, 24 GB)Q4_K - Medium64.7 tok/s916.5 tok/sllamafileref
M3 (GPU count not published, 16 GB)Q4_K - Medium61.6 tok/s229.7 tok/sllamafileref
M1 Ultra (GPU count not published, 128 GB)Q4_K - Medium57.1 tok/s283.8 tok/sllamafileref
M2 (10-core GPU, 8 GB)Q4_K - Medium56.5 tok/s804.8 tok/sllamafileref
M2 (10-core GPU, 16 GB)Q4_K - Medium55.6 tok/s801.8 tok/sllamafileref
M2 (10-core GPU, 24 GB)Q4_K - Medium54.8 tok/s813.7 tok/sllamafileref
M1 (8-core GPU, 8 GB)Q4_K - Medium40.4 tok/s607.4 tok/sllamafileref
M1 (8-core GPU, 16 GB)Q4_K - Medium40.2 tok/s585.7 tok/sllamafileref
M1 (7-core GPU, 8 GB)Q4_K - Medium38.5 tok/s533.0 tok/sllamafileref
M1 (7-core GPU, 16 GB)Q4_K - Medium37.9 tok/s530.9 tok/sllamafileref
M2 (8-core GPU, 16 GB)Q4_K - Medium35.3 tok/s523.2 tok/sllamafileref
M2 (8-core GPU, 8 GB)Q4_K - Medium34.5 tok/s523.2 tok/sllamafileref
M5 Max (32-core GPU, 36 GB)M4 Max (40-core GPU, 128 GB)M4 Max (40-core GPU, 64 GB)M4 Max (40-core GPU, 48 GB)M3 Ultra (80-core GPU, 512 GB)M3 Ultra (80-core GPU, 256 GB)M2 Ultra (60-core GPU, 128 GB)M2 Ultra (60-core GPU, 64 GB)M2 Ultra (60-core GPU, 192 GB)M4 Max (32-core GPU, 36 GB)M4 Max (GPU count not published, 128 GB)M2 Max (38-core GPU, 32 GB)M1 Ultra (64-core GPU, 128 GB)M3 Max (40-core GPU, 48 GB)M3 Max (40-core GPU, 128 GB)M1 Ultra (48-core GPU, 128 GB)M3 Max (30-core GPU, 36 GB)M3 Max (30-core GPU, 96 GB)M2 Max (30-core GPU, 32 GB)M1 Max (32-core GPU, 32 GB)M1 Max (32-core GPU, 64 GB)M2 Ultra (GPU count not published, 128 GB)M4 Pro (20-core GPU, 24 GB)M4 Pro (20-core GPU, 48 GB)M4 Pro (20-core GPU, 64 GB)M4 Pro (16-core GPU, 64 GB)M4 Pro (16-core GPU, 48 GB)M4 Pro (16-core GPU, 24 GB)M3 Max (40-core GPU, 64 GB)M1 Max (24-core GPU, 64 GB)M2 Pro (19-core GPU, 32 GB)M2 Pro (19-core GPU, 16 GB)M5 (10-core GPU, 32 GB)M5 (10-core GPU, 16 GB)M1 Max (24-core GPU, 32 GB)M2 Pro (16-core GPU, 32 GB)M2 Pro (16-core GPU, 16 GB)M3 Pro (18-core GPU, 36 GB)M3 Pro (14-core GPU, 36 GB)M3 Pro (14-core GPU, 18 GB)M3 Pro (18-core GPU, 18 GB)M1 Pro (16-core GPU, 16 GB)M1 Pro (16-core GPU, 32 GB)M4 (10-core GPU, 16 GB)M4 (10-core GPU, 32 GB)M4 (10-core GPU, 24 GB)M1 Pro (14-core GPU, 16 GB)M1 Pro (14-core GPU, 32 GB)M4 (GPU count not published, 16 GB)M3 (10-core GPU, 16 GB)M4 (8-core GPU, 16 GB)M3 (10-core GPU, 24 GB)M3 (GPU count not published, 16 GB)M1 Ultra (GPU count not published, 128 GB)M2 (10-core GPU, 8 GB)M2 (10-core GPU, 16 GB)M2 (10-core GPU, 24 GB)M1 (8-core GPU, 8 GB)M1 (8-core GPU, 16 GB)M1 (7-core GPU, 8 GB)M1 (7-core GPU, 16 GB)M2 (8-core GPU, 16 GB)M2 (8-core GPU, 8 GB)

benchmarks.json — full dataset  ·  models.json — model summaries  ·  benchmarks.csv — CSV export

See all models →