Skip to the content.

idea-scout triage — progress advantage: a free RL-post-training step reward (issue #1122)

Closes the daily idea-scout candidate #1122 (tools/idea_scout.py, filed 2026-06-28). The scout judges whether a candidate is new and on-topic; this note is the human triage it hands off — adopt as a capability, defend against as a threat, or cite as prior art (see docs/idea-scout.md). Verdict: prior art to cite — at the thesis level, with a sharp boundary. The paper’s “the existing pipeline already contains the step-level signal, so you need no dedicated trained judge” is the reward-modeling-side sibling of fak’s witness referee (dos verify grades the recorded diff, not a trained reward model). But its signal is the policy’s OWN log-probability — a model-internal, confidence-like score — and fak’s recorded position is detector-is-not-the-floor / witness-not-self-report: a self-report read off the model’s own probabilities hosts ABOVE fak’s trust floor, never on it. Not adopted as a capability (fak does not RL-post-train, and has no paired reference policy for its served models, so the “free lunch” is unavailable on fak’s actual fleet; no kernel surface) and not a threat. One recorded residual — progress advantage as a cheap PROPOSAL signal for fak’s trajectory graders — is NOT adopted. No code change.

Source: https://arxiv.org/abs/2606.26080 — “Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents”, Changdae Oh, Wendi Li, Seongheon Park, Samuel Yeh, Tanwi Mallick, Sharon Li (submitted 2026-06-24). Read from the arXiv abstract as surfaced by the scout on 2026-06-28; this is a surface read of the abstract, not a paper audit.

What it is

A reward-modeling / RL-post-training result, not a serving system, a protocol, an attack, or a defense. The premise: process reward models (PRMs) give fine-grained, step-level evaluation of an LLM’s reasoning, but building one for an agentic setting is prohibitively hard — long-horizon interactions, irreversible actions, and stochastic environment feedback defeat both human annotation and Monte-Carlo estimation at scale.

The claim is that RL post-training already contains the ingredients for step-level scoring, so a dedicated PRM need not be trained at all. Under a general stochastic MDP, the paper derives that the log-probability ratio between the RL-trained policy and its reference policy exactly recovers the optimal advantage function — a signal it names progress advantage. It is therefore annotation-free, domain-agnostic, and available as a byproduct of the standard RL post-training pipeline (the same implicit- reward shape the DPO/PPO-reference family uses: r(s,a) ∝ log[ π_RL(a|s) / π_ref(a|s) ]).

The paper validates the signal on three applications — test-time scaling, uncertainty quantification, and failure attribution — across five benchmarks and four model families, reporting that it consistently beats confidence-based baselines and, with no task-specific training, surpasses dedicated trained reward models.

It proposes a derivation + a measurement, used at inference to score an existing agent’s steps. It does not propose a kernel component, a runtime mechanism, an MCP server, or a serving primitive fak would implement.

The three triage questions

fak is an agent kernel: one Go binary at the tool-call seam that adjudicates every tool call before it runs — a default-deny capability floor (security gate) plus do-the-shared-setup-once cross-turn reuse (performance gate). Against that mission:

Why this is prior art to cite — a free pipeline byproduct beats a trained judge

fak’s whole referee design is a refusal to train or trust a dedicated judge. The dos verify / dos commit-audit truth syscall grades a claimed result against the diff git actually recorded — a derived, annotation-free witness already produced by the normal commit pipeline — never against a learned reward model and never against the commit message the author wrote (a forgeable self-report). The recorded economics is exactly the paper’s headline: the existing pipeline already contains the step-level signal you would otherwise pay to train a judge for. The paper makes that argument from the RL-reward-modeling direction (the post-training run already encodes the optimal advantage in the policy/reference ratio); fak makes it from the evidence direction (the version-control run already encodes the step’s effect in the diff). Both replace an annotation-hungry trained judge with a free byproduct of the pipeline that was going to run anyway — and the paper’s empirical result that the free signal outperforms dedicated trained reward models is independent support for that “don’t build the judge” bet.

The sharp boundary — the signal is a self-report, and fak keeps those off the floor

The citation stops at the economics. What the two signals are is opposite. Progress advantage is the policy’s own log-probability — a model-internal, self-referential score (the paper benchmarks it precisely against confidence-based baselines, i.e. it lives in the same family as the model’s own confidence). fak’s recorded position is detector-is-not-the-floor / witness-not-self-report (CLAIMS.md “Honest ceiling”: detection is ~100% evadable, non-load-bearing; cf. defensive-misdirection and AUC-is-not-detection). A score read off the model’s own probabilities is exactly the kind of self-report fak hosts above its load-bearing floor, never on it.

So the honest fence is the same one drawn for the CoT-as-explanation sibling (#1008) and the contested attention-as-explanation demotion in fak’s reward-over-spans epic (#861/#862): a self-generated signal — CoT text, attention mass, or a policy log-prob ratio — is a proposal to be verified against an action-level witness, never an accepted causal account. fak’s witness is external, action-level, and git-recorded — a fact the model cannot author; progress advantage is internal confidence the model does author. They sit on opposite sides of the trust boundary and must be cross-linked, not conflated.

The one recorded residual — a proposal signal for the trajectory graders, NOT adopted

There is a genuine, fak-shaped follow-on, recorded here precisely so it is not mistaken for a shipped lever. fak’s dispatch fleet and dojo gym already grade agent trajectories step by step — closure honesty + the witness ledger on the fleet side, and the fak dojo predict→run→measure→eval→calibrate loop (#951) on the gym side. The paper’s failure-attribution application — which step of a failed trajectory lost the most progress? — maps directly onto a real fak need (attributing a failed dispatch to the step that broke it, today done by the diff/witness referee, not a per-step score). Progress advantage is a candidate cheap proposal signal to rank steps for that attribution.

It is not adopted in this increment, for two honest reasons:

  1. It is unavailable on fak’s models. Computing it needs paired RL-policy and reference-policy log-probs for the served model; fak’s fleet is third-party / local weights with no paired reference policy, so the signal cannot be produced where the need is. (This is the same “the prerequisite pipeline isn’t fak’s” wall as the adopt question above.)
  2. It can only ever be a proposal, never the floor. Even where computable, a model-internal score is, by fak’s own thesis, hosted above the evidence floor — it may propose which step to look at, but the attribution that counts is still the action-level, git-recorded witness (dos verify). A trajectory grader that decided on the log-prob ratio would be trusting a self-report.

So the residual is filed as a candidate proposal signal for the dispatch/dojo trajectory graders — gated behind a paired-policy availability fak lacks, and explicitly fenced as never-the-floor. Not a feature, and explicitly not shipped here.

Triage decision

Action: this note is the recorded triage; close #1122 as triaged → prior art recorded, not adopted as a capability or defended as a threat. No code change in this increment: the right artifact for a reward-modeling-result candidate is the recorded verdict and the named boundary, not a speculative feature.

Scout calibration (no code change). The candidate surfaced under topic agent-model-arch (score 44) on a genuine training (title) term plus a freshness bonus (≤30d, 3d old) — correctly on-topic. It is the third recent agent-model-arch hit that is a training-time result fak cannot adopt as a capability (after the #1008 CoT-training triage; cf. the #861/#862 reward-over-spans line) — an on-topic-by-keyword / off-mission-by- content pattern for the training scorer worth watching. It is not a scoring bug: the scout judges new and on-topic, never worth building, and these training papers still yield citable thesis-level prior art (this one included), so suppressing the training term would lose signal, not just noise. The scout behaved as designed and handed the call to human triage (see docs/idea-scout.md); there is no change to tools/idea_scout.py — only this recorded verdict.