Skip to the content.

Never measured — the comparisons that would actually decide it

Each row is a facet a buyer materially weights where fak has no head-to-head against the thing that buyer would otherwise use. These are not gaps in fak; they are gaps in the evidence, and they are the highest-value benchmark work in the repo — every one of them is currently deciding a purchase by default.

weight use case × facet why it is missing next
0.35 solo-max × session-longevity the deciding comparison for this buyer – fak’s compaction versus Claude Code’s own native auto-compaction, over a session long enough for both to fire repeatedly – has never been run. The repo has fak’s per-fire shed (~107K tokens against a 48K budget) but no measurement of the alternative on the same session, and docs/benchmarks/AUTHORITY-GENERATED-SAMPLE.md:465 states plainly that sessionbench has no checkpoint/resume. A number for one arm is not a comparison. run one long real task twice (bare claude vs fak manage -- claude), same prompt and same tools, and report turns-to-context-exhaustion and re-explanation events for both arms
0.10 platform-team × portability this buyer’s portability question is provider and agent lock-in, not the model x backend grid; the committed support matrix answers a different question and there is no comparison against what a plain proxy locks in. score both against the same swap list (provider, model, agent, host) with a CI witness per swap point
0.10 researcher × injection-control the formal-isolation tier costs more than this buyer’s whole tolerance, so it is not their next-best option; the honest alternative is their own ad-hoc filtering, which nobody has measured. add an unguarded-harness arm to the AgentDojo run so the comparison is against real practice
0.05 solo-max × injection-control the formal-isolation tier fak’s 0/38 AgentDojo result is at parity with (CaMeL, MELON) costs more to deploy than this buyer’s entire tolerance, so it is not their next-best option – and no ASR measurement exists against what they would actually use, which is the host agent’s own permission prompts. run AgentDojo against bare claude with default permission prompting as the alternative arm, so the comparison is against the thing this buyer actually has
0.05 solo-max × observability the alternative here is the agent’s own transcript, not an OTel tracing stack, and nobody has decomposed what a bare Claude Code transcript does and does not let you reconstruct after the fact. take one failed real session and score both artifacts against the same five decision classes (model traffic, cache reuse, compaction, tool verdicts, recovery)
0.05 solo-max × portability for this buyer portability means keeping their agent and their login, not the model x backend grid; there is no committed comparison of what bare Claude Code locks in versus what fak locks in. enumerate the swap points (agent, provider, model, host) for both and mark which are actually exercised in CI
0.05 regulated × portability no committed comparison of vendor lock-in against a container-plus-policy incumbent; the model x backend grid is not what this buyer means by portability. enumerate the exit path from each and mark which steps have a witness
0.05 local-first × injection-control the formal-isolation tier costs more to deploy than this buyer’s whole 12h tolerance, so it cannot be their next-best option, and nobody has measured ASR against what they would actually use. run AgentDojo with the local agent’s own approval prompts as the alternative arm
0.05 researcher × portability no committed comparison of what an SDK-plus-logging setup locks in versus what fak locks in. score both on the same swap list with a witness per swap point