Skip to the content.

Study: Caveman — shape-specific context transforms with measured fallback

Observed: 2026-08-13
Source event: upstream HEAD c72984e4392c7a154e55c11dbf445f01ce5c35d4 committed 2026-08-13T17:27:59+02:00
Source state: public GitHub repository, default branch main, exact pinned revision
Platform context: static source study on Windows; no Caveman binary was executed and no upstream performance claim is adopted as a fak claim
Refresh trigger: upstream engine license/change date, a material engine/compressors, engine/pixel, or engine/ccr release, or closure of filed issues #6668–#6670

Value frame

Problem centrality is Enabling. P1 managed context is direct; P2 net-true efficiency requires tuned-baseline ablation; P3 adaptation stays bounded through typed gates and raw fallback; P4 operations require visible decisions and refusal reasons. The smallest spine for each surviving borrow is therefore an offline or one-shape witnessed gate, not a broad compression engine.

What Caveman is optimizing

Caveman targets developers whose coding agents repeatedly carry large tool outputs and tool schemas. Its v2 engine treats compression as a portfolio of shape-specific transforms rather than one summarizer: structured data, search results, terminal/log output, code, and dense text each receive a separate applicability decision. Tests emphasize adversarial retention, round-trip behavior, explicit thresholds, and fallbacks. The repository also exposes memory, cache, proxy, browsing, MCP, UI, and packaging surfaces around that engine.

That worldview differs from fak’s primary one in degree, not kind. Caveman optimizes immediate provider-token reduction across popular harnesses; fak optimizes a kernel-owned, auditable context lifecycle where cache survival, reversibility, and net value are first-class. Borrowing therefore means adding measured transforms at fak’s existing ownership seams, not importing Caveman’s engine or its claims.

Coverage and completeness critic

The study read the README and honest-number/benchmark docs as maps, then inspected implementation and tests across:

No load-bearing top-level subsystem was left unopened. Generated docs/site assets, font binaries, screenshots, and distribution archives were inventoried but not byte-reviewed because they cannot change the borrow decision; their licenses/notices were checked. Open issues and PRs were treated as direction, never shipped proof.

License and provenance gate

The repository is mixed-license. Root/open interface surfaces include MIT-licensed portions, while the engine and several runtime packages use Business Source License 1.1 with an Apache-2.0 change date of 2028-08-11; pixel assets carry additional font notices. Public visibility is not permission. All three filed borrows are therefore INSPIRE-ONLY: independently implement behavior against fak-owned fixtures, copy no expressive code/tests/comments/assets, and retain the pinned source only as prior-art provenance.

Candidate cards

Borrow Source anchor Axis Their-worldview reason fak witness on-axis Route Filed
Shape-specific terminal/search/log result codec with raw fallback engine/compressors/searchresult.go:28-67@c72984e; terminal.go:26-66@c72984e; log.go:22-71@c72984e Resident tool-result tokens while preserving anchors and deterministic recovery Coding agents repeatedly carry semi-structured command output; format knowledge saves more than generic truncation PARTIAL: internal/ctxmmu/mmu.go, toolpages.go, and capbody.go own budgets/CAS/paging but have no shape codec; TOON #3064 owns tabular JSON only INSPIRE-ONLY #6670
Provider-aware text-to-image density gate engine/pixel/applicability.go:40-112@c72984e; density.go:16-91@c72984e; transform_openai.go@c72984e Choose text versus image representation only when provider-token value is positive Dense monospace output can be cheaper as vision input, but only for supported models/content/geometry ABSENT: fak has image geometry, ingestion, screenshot dedup, and VLM epic #4033, but no outbound tool-result renderer or density decision INSPIRE-ONLY #6668
Annotation-aware minimization for hot tool schemas engine/compressors/toolschema.go:26-118@c72984e; toolschema_annotations.go:18-106@c72984e Remove only prose structurally redundant with type/requiredness while retaining semantic constraints Harnesses resend schema prose every turn; blanket deletion is unsafe, so redundancy is field-local and fail-closed PARTIAL: #3229/#3231/#3232 defer cold schemas, but the remaining hot schemas have no annotation-aware minimizer INSPIRE-ONLY #6669
TOON/tabular JSON encoding engine/compressors/toon.go@c72984e; json_strategy.go:14-47@c72984e Reversible tabular JSON token density Repeated objects are cheaper in a compact tabular wire form PRESENT-on-axis / already tracked: internal/toon and epic #3064 with open scorecard #3068 and governed lossy follow-on #3343 INSPIRE-ONLY; do not refile #3064
Cold tool-schema deferral Caveman tool-schema compression neighborhood; fak comparison at internal/gateway tool-search seams Avoid sending unused tool definitions at all Schema tokens are an always-sent floor PRESENT-on-axis: shipped #3231/#3232 under epic #3229 defer cold schemas; this dominates minification for cold tools Stay with fak design #3229
Exact response cache engine/ccr/store.go:10-41@c72984e; store_sqlite.go:90-180@c72984e Return a stored completion for an exact repeated request Repeated deterministic requests should avoid provider work entirely DIVERGENT: fak deliberately centers provider prefix/KV reuse plus deterministic replay/audit rather than silently substituting a prior model response. Exact-response substitution changes freshness and stochastic semantics; no leaf filed without a named safe workload Note only
Cross-session text memory mem/README.md@c72984e and mem/*.go@c72984e Persist/retrieve working facts across sessions Coding users want continuity without replaying full transcripts PRESENT-on-axis: fak memory/recall core images, internal/memq, and session images already provide governed persistence and retrieval No borrow
Generic token-level compressor compressor/eval harness and engine/evals/harness.go@c72984e Semantic compression under quality thresholds Different content needs measured quality, not one ratio PRESENT/on backlog: Compressor work and closed #3204 already establish the plugin/eval direction; no Caveman-specific leaf survives No duplicate #3204

Negative knowledge and direction

Filed backlog

Each issue carries both the pinned upstream anchor and the fak seam, the on-axis verdict, the BSL-driven INSPIRE-ONLY route, a first checkable step, scope guard, and witness. The issues were deduplicated against all open/closed fak issues and received the derived class:dev label through tools/issue_lane_router.py.

Companions

Addendum — opt-in user profiles, not silent defaults (2026-08-13)

A user follow-up identified a separate product requirement: some operators want the skill shape itself, not only lower internal context cost. fak should expose that value, but the two referenced systems affect different axes and must not be collapsed into one switch.

Existing fak seam

Current trunk already ships the safe base mechanism:

That makes Caveman response shape PARTIAL-on-axis, not absent. The missing piece is a named, intensity-bearing profile with reproducible metadata and a benchmarked quality/cost envelope.

Two orthogonal feature flags

Concern Safe selector Default May change Must never change Filed
Caveman-like response shape --output-style caveman:medium (later low|high|auto) default sentence/bullet shape, connective prose, response verbosity facts, qualifiers, safety, code, commands, diagnostics, identifiers, explicit user format #6701
Ponytail-like implementation restraint --work-profile ponytail standard planning preference toward no-op/deletion/config/native/stdlib and smallest correct diff requested scope, correctness, tests, security, authorization, compatibility, migration, proof #6700

The flags compose explicitly and never imply each other. --work-profile ponytail --output-style caveman:medium requests both; selecting only one leaves the other axis unchanged. Precedence is structural: system policy and explicit user requirements > repository instructions > work profile > output style. Unknown names fail closed. Every active profile records its name, pinned inspiration revision, injected-fragment digest, selection reason, and exact disable command in session metadata.

Ponytail was checked at DietrichGebert/ponytail@2ed6c52c9d7e5e56942508591085fd45dea277d3. Its operational mechanism is the simplicity ladder in skills/ponytail/SKILL.md:13-92 with explicit safety carve-outs at lines 115-124. Unlike Caveman response shape, this changes implementation decisions, so treating it as terse would be unsafe and semantically wrong. Ponytail is MIT at the pin; #6700 uses an attributed fak-native ADAPT route without the persona or always-on hook behavior.

Superset claim gate

The architecture can make fak’s selectable envelope a strict functional superset, but the phrase “without question a superset” remains a gated claim, not current marketing. It becomes supportable only when:

  1. #6701 and #6700 ship through the governed syspromptmmu seam with default-off/fail-closed tests, precedence tests, preservation/safety fixtures, captured context plans, and visible disable paths;
  2. benchmark epic #6674, especially Caveman ablation #6683 and Ponytail correctness/robustness tracks #6687-#6691, proves task quality and net-true cost against pinned comparator configurations; and
  3. the compatibility matrix distinguishes behavioral coverage, safe composition, and measured advantage rather than using one undifferentiated “compatible” badge.

Until those witnesses exist, the honest wording is: fak has the governed substrate and filed opt-in profiles needed to cover both user experiences; comparative parity is not yet proven.

Implemented spine — family, implementation, and intensity are separate (2026-08-13)

The first native spine now treats a response profile as three dimensions rather than one overloaded style word:

family : implementation : intensity

The accepted Caveman-compatible native spellings are caveman:native:low, caveman:native:medium, and caveman:native:high. Existing concise|brief|terse|minimal and native:low|medium|high remain aliases for fak’s native scale. The fak agent --output-style flag drives the owned system block through FAK_OUTPUT_STYLE, and the captured SystemBlock records the canonical style, family, and intensity while leaving the resident cache prefix byte-identical.

This grammar leaves a deliberate slot for an eventual caveman:original:<intensity> adapter without pretending fak-authored bytes are upstream-original. original, caveman:original, unsupported intensities, and unrelated families currently fail closed. An original adapter must pin and attribute the source revision, preserve the upstream license, expose the exact fragment digest, and pass the same safety/precedence tests before it is admitted.

Mix-and-match remains orthogonal: response shape composes later with --work-profile ponytail:{native|original}:{low|medium|high} rather than making Ponytail an output style. The current commit ships only the response-profile spine; original-source adapters, Ponytail work profiles, auto, guard/harness propagation, and a general repeated --profile expression remain filed follow-ons rather than silently deferred behavior.