Skip to the content.

The fak caching ladder

What this is. The same subject — how fak caches so a long, stop-start agent session stays cheap — written five times, once for each kind of reader. Each rung is pitched at both a human audience and a model/agent capability tier: a newcomer or a small 8B agent is served at Level 1; a frontier / fable-tier reader is served at Level 5. Pick your rung, or climb the whole ladder.

Caching in fak spans two very different layers — the provider’s prompt cache that fak steers on the API wire, and the kernel-owned KV cache fak owns when it runs the model itself. A single flat page either loses the beginner or bores the expert. This ladder splits the difference: each rung assumes exactly what the one below it taught.

The five rungs

Level Who it’s for What you’ll get
1 · What is caching? A newcomer, or a small/8B agent that needs the one-screen version. No jargon. What a prompt cache is, why it matters, what to type, how to tell it helped.
2 · Managed cache in practice A competent dev or mid-tier agent running fak manage -- claude. --managed-cache on\|off\|auto, the defaults, Pro/Max headroom, how to verify.
3 · Cache economics & the wire A senior engineer / capable agent who wants the mechanism and the money. Sliding window vs 1h tier, the read/write multipliers, byte-exact prefix, provider matrix.
4 · The kernel-owned KV cache A platform engineer / strong agent running the model in-kernel. Addressable KV cache, bit-exact span eviction, prefix cloning, WITNESSED vs OBSERVED accounting.
5 · The caching frontier A fable-tier reader / researcher who wants the SOTA and the open problems. vCache enablement, SOTA optimizations, cross-provider futures, the honest open edges.

How to read it

Each rung ends with a nav footer linking its neighbours and back here.

See also