Skip to the content.

Qwen3.8-27B M3 Pro evidence (2026-08-20)

For the current cross-hardware view, use the canonical Qwen performance index. This page retains the dated M3 Pro envelope and its detailed witnesses.

Generated current readout

Latest retained Metal receipt: native fak Metal ran Qwen3.8-27B Q4_K_M at 2.3–2.9 decode tok/s and 3.2–8.4 full-prefill tok/s on the 18-GPU-core, 36 GiB Apple M3 Pro. This was an accepted result for its dated envelope, not parity with llama.cpp. Its 2026-08-27 review window passed without a comparable renewal, so it is now historical and is reaped from the active front-door presentation.

Source of truth: metal-native-run-summary.json, observed 2026-08-20 and reviewed through 2026-08-27. No newer comparable, accepted Metal full-run receipt supersedes it. The exact model is the 17,106,775,008-byte unsloth/Qwen3.8-27B-GGUF Q4_K_M artifact at revision f1bfb127c64f7072bdd2cad55f258b9c8b2910fe, SHA-256 7e78da5d7e3ae28d178121f58646953305f3e5bd3cb46f4a75584e8b6c6fe169.

Current numbers

Measure Result Scope
Decode 2.3 tok/s text after full prefill
Decode 2.9 tok/s JSON after full prefill
Fully cached decode 0.4–1.3 tok/s observed range; do not combine with the full-prefill rows
Full prefill 3.2 / 5.9 / 8.4 tok/s text / JSON / tool probes; prompt lengths differ
Time to ready 103.858 s cold / 34.897 s warm-file-cache native Metal service startup
Max RSS 18.92–19.56 GB two captured runs
macOS peak footprint 42.54–45.63 GB OS footprint is not RSS; zero swaps recorded
Functional acceptance PASS exact text, strict JSON, and admitted tool probes

The same witness records 0.03 tok/s before the decode fix, so the accepted path is roughly 77–97× faster than that broken baseline. That ratio does not establish parity with another engine.

What changed relative to Qwen3.6

The premise that native fak Qwen3.6 had reached near parity with llama.cpp on this Mac is not supported by the accepted page. The Qwen3.6 parity bar records:

Qwen3.6-27B Q4_K_M path Prefill Decode
llama.cpp Metal b9707 51.55 tok/s 7.29 tok/s
fak resident-Q4_K Metal 2.6 tok/s at P=27; 7.3 at P=940 1.2 tok/s

Against fak’s last accepted Qwen3.6 Metal decode, Qwen3.8’s 2.3–2.9 tok/s is 1.92–2.42× higher. Against the older Qwen3.6 llama.cpp bar, it is only 32–40% as fast. These are directional comparisons, not model-to-model parity: the generation, prompt/corpus, code revision, and measurement date differ, and there is no accepted same-artifact Qwen3.8 llama.cpp row yet.

The main delta is implementation maturity, not a demonstrated architectural windfall:

August 27–28 update

The canonical Qwen3.8 artifact pin, native-receipt/readmit lineage tests, Vulkan GDN route and decode, and host-visible Vulkan Q4_K staging have landed. The dated trajectory dogfood and usage-outcome snapshot record supporting operational evidence. These are implementation, support, and diagnostic facts—not a newer accepted speed result. Metal and AMD/Vulkan remain awaiting comparable, quality-complete full-model remeasurement.

Evidence status and next comparison

For benchmark contract requirements, see BENCHMARK-CONTRACT-MAP.md.