Skip to the content.

GLM-5.2: the full DSA forward on the pure fak kernel, witnessed on the sm_80 GPU server (2026-06-23)

Goal: prove GLM-5.2 runs on fak’s own kernel on the real datacenter GPU server — pure fak, not a third-party engine. This note records the on-device witness (grounded in the node’s own go test -tags cuda exit code and run log, not self-report) and is explicit about the boundary: the kernel-math forward is proven on real sm_80 silicon; full-size 753B serving still routes through llama.cpp (the labeled comparison), because fak’s native engine is f32 / no quantized-GGUF device GEMM today.

Superseded by progress (#917; see the staged plan). The “753B serving still routes through llama.cpp … no quantized-GGUF device GEMM” boundary above was the 2026-06-23 snapshot; the quantized device GEMM + --cpu-offload-experts rungs have since landed and fak’s own engine loads the full 466 GB model natively (2026-06-25 native-serve note). The sm_80 DSA-forward witness this note records stands; the serving boundary does not.

The headline: the full DSA forward on sm_80, cosine = 1.000000

Until today the on-hardware GLM-5.2 pure-kernel captures were split across two arches:

This run closes that gap. The complete GLM-MoE-DSA forward — MoE/FFN experts + router + vocab head, the DSA attention dense projections on k_q8_gemm, and the DSA sparse-attention compute itself on k_dsa_sparse_attend — now runs on fak’s CUDA kernel on a real sm_80 datacenter GPU, argmax-exact, cosine = 1.000000 vs the all-host CPU Q8 forward. The node’s own run log (fresh clone of origin/mainnvcc -arch=sm_80 build → isolated -tags cuda test):

=== HEAD 44aa3b6 ===
=== build libfakcuda.a (sm_80) ===
[cuda] nvcc compile kernels (sm_80) ...
[cuda] go build -tags cuda ./internal/compute/ ... OK build
=== go test -tags cuda -run TestCUDAGLMMoeDsaBackendForward ./internal/model/ -v ===
=== RUN   TestCUDAGLMMoeDsaBackendForward
    glm_dsa_cuda_test.go:63: GLM-MoE-DSA forward with MoE/FFN+head + DSA attention
      projections (k_q8_gemm) + DSA sparse attention (k_dsa_sparse_attend) on cuda
      backend: cosine=1.000000 argmax cpu=40 cuda=40 tier=sm_80 class=approx
--- PASS: TestCUDAGLMMoeDsaBackendForward (0.16s)
PASS
ok  github.com/anthony-chaudhary/fak/internal/model   0.853s
=== GLM GPU WITNESS DONE rc=0 ===

The hard compute.DSASparseBackend type-assert inside the test fails the test rather than silently falling back to the host loop — so its PASS confirms the sparse attention executed on the device kernel, not host. Reproduce on any sm_80+ CUDA node:

bash private GPU witness runner    # clone origin/main -> nvcc -arch=sm_80 -> the isolated witness

The honest boundary: this is kernel-math, not 753B serving

The witness runs a tiny GLM-DSA fixture (kernel correctness), not the 753B weights. So what is proven is precise and bounded: fak’s own kernel computes GLM-5.2’s DSA forward bit-faithfully on real sm_80 silicon. It is not “fak serves the full 753B.” fak’s native engine is f32 / has no quantized-GGUF device GEMM / no multi-GPU NCCL today, so it cannot load the 753B Q4 checkpoint — that is the labeled gap, not a claim.

The comparison baseline (llama.cpp), and why it was brought down

Full-size GLM-5.2 does serve on the same sm_80 server today — via llama.cpp with CPU offload (-ngl 99 --n-cpu-moe 99): a community Q4_K_M GGUF (≈424 GB) resident in host RAM with the MoE experts on CPU and the dense/attention layers on the GPUs, answering on an OpenAI-compatible port at ~2.66 tok/s decode (~5.1 tok/s prompt). This confirms the CPU-offload path makes the full 753B serveable on sm_80 — no Hopper-class DSA kernel required, exactly as expected — but the engine is llama.cpp, not fak, so it is recorded here strictly as the comparison baseline. (Like other reasoning models on a tight token budget, it tends to spend the budget on reasoning_content before the literal answer — a harness caveat, not a serving failure.)

Provenance audit (so the comparison server is never mistaken for a fak dependency): it was a manually-started, detached process — parent = init(1), with no systemd/cron/supervisor, so it does not auto-restart — serving the community GGUF from a local llama.cpp build. No fak process referenced it. To give the pure-fak witness a fully idle machine, it was brought down cleanly (all GPUs returned to 0 MiB used); the exact one-line relaunch command was recorded first, so the comparison is restartable on demand.

What is proven vs not (labeled)

See also GLM52-DSA-SPARSE-ATTENTION-ON-PURE-KERNEL-2026-06-23.md (the sparse-attention seam + the sm_89 capture) and GLM52-DSA-PROJECTIONS-ON-PURE-KERNEL-GPU-SERVER-2026-06-22.md (the dense-projection slice + the original sm_80 dense capture).