Never measured — the comparisons that would actually decide it
Each row is a facet a buyer materially weights where fak has no head-to-head against the thing that buyer would otherwise use. These are not gaps in fak; they are gaps in the evidence, and they are the highest-value benchmark work in the repo — every one of them is currently deciding a purchase by default.
| weight | use case × facet | why it is missing | next |
|---|---|---|---|
| 0.35 | solo-max × session-longevity | the deciding comparison for this buyer – fak’s compaction versus Claude Code’s own native auto-compaction, over a session long enough for both to fire repeatedly – has never been run. The repo has fak’s per-fire shed (~107K tokens against a 48K budget) but no measurement of the alternative on the same session, and docs/benchmarks/AUTHORITY-GENERATED-SAMPLE.md:465 states plainly that sessionbench has no checkpoint/resume. A number for one arm is not a comparison. |
run one long real task twice (bare claude vs fak manage -- claude), same prompt and same tools, and report turns-to-context-exhaustion and re-explanation events for both arms |
| 0.10 | platform-team × portability | this buyer’s portability question is provider and agent lock-in, not the model x backend grid; the committed support matrix answers a different question and there is no comparison against what a plain proxy locks in. | score both against the same swap list (provider, model, agent, host) with a CI witness per swap point |
| 0.10 | researcher × injection-control | the formal-isolation tier costs more than this buyer’s whole tolerance, so it is not their next-best option; the honest alternative is their own ad-hoc filtering, which nobody has measured. | add an unguarded-harness arm to the AgentDojo run so the comparison is against real practice |
| 0.05 | solo-max × injection-control | the formal-isolation tier fak’s 0/38 AgentDojo result is at parity with (CaMeL, MELON) costs more to deploy than this buyer’s entire tolerance, so it is not their next-best option – and no ASR measurement exists against what they would actually use, which is the host agent’s own permission prompts. | run AgentDojo against bare claude with default permission prompting as the alternative arm, so the comparison is against the thing this buyer actually has |
| 0.05 | solo-max × observability | the alternative here is the agent’s own transcript, not an OTel tracing stack, and nobody has decomposed what a bare Claude Code transcript does and does not let you reconstruct after the fact. | take one failed real session and score both artifacts against the same five decision classes (model traffic, cache reuse, compaction, tool verdicts, recovery) |
| 0.05 | solo-max × portability | for this buyer portability means keeping their agent and their login, not the model x backend grid; there is no committed comparison of what bare Claude Code locks in versus what fak locks in. | enumerate the swap points (agent, provider, model, host) for both and mark which are actually exercised in CI |
| 0.05 | regulated × portability | no committed comparison of vendor lock-in against a container-plus-policy incumbent; the model x backend grid is not what this buyer means by portability. | enumerate the exit path from each and mark which steps have a witness |
| 0.05 | local-first × injection-control | the formal-isolation tier costs more to deploy than this buyer’s whole 12h tolerance, so it cannot be their next-best option, and nobody has measured ASR against what they would actually use. | run AgentDojo with the local agent’s own approval prompts as the alternative arm |
| 0.05 | researcher × portability | no committed comparison of what an SDK-plus-logging setup locks in versus what fak locks in. | score both on the same swap list with a witness per swap point |