Part of my research on robust evaluation of adaptive systems. Branch closed 2026-09-29.
AI Evaluation · Agent Reliability · Experimental Systems
Evaluator Assurance: What Survived
Published Sep 29, 2026I tested seven plausible ways to make evaluator-driven AI systems more trustworthy across AgentRewardBench and τ-bench. Most weakened or failed when exposed to preregistration, held-out data, operational metrics, and cheap baselines. The useful result was narrower than the architecture I started with.
This is a falsification study, not a framework proposal. The research value comes from determining what did not earn authority, and preserving enough evidence that a skeptical reader can inspect why.
3,086
trajectory records across two studies
5,530
paid repeat judgments (R=5)
$46.18
measured inference spend
7
mechanisms frozen and tested
13
documented process errors
4-day
intensive research sprint
The question
LLM-backed systems increasingly use evaluators to decide whether an agent succeeded, whether to retry, whether to escalate, or whether to trust a generated result. When those evaluators make mistakes, the intuitive response is often: sample more, add another judge, add routing, add deterministic guards, combine everything.
The project asks a more precise question:
Which of those mechanisms actually earn their cost and authority when tested against real trajectories, with preregistered gates and a held-out corpus they have never seen?
Corpora
| Corpus | n | Evidence class |
|---|---|---|
| AgentRewardBench rev b6d17e6 | 1,106 | Preregistered |
| τ-bench historical_trajectories/ | 1,980 | Fresh / never touched |
Mechanisms tested
- M1Repeated execution of the same judge
- M2Alternate-evaluator pairing / escalation
- M3–M6Deterministic pre-check rules (R1–R4, held-out)
- M7Static required-conjunct matching (RC1, fresh corpus)
Result 1 — Preregistered
Repetition had almost nothing to average over
gpt-4o-2024-11-20 · temperature 0.0 · seed 0 · judge stage only
Five repeated calls were made on each of 1,106 AgentRewardBench trajectories — 5,530 calls total. If repeated sampling were recovering meaningful variance, a substantial fraction of cases would land in the middle of the k=0–5 distribution.

The deployed judge was not perfectly deterministic. It was simply stable enough that repeated sampling had almost no useful error to recover. Only 20 of 1,106 cases showed any intermediate correctness across five draws.
20
of 1,106 cases had k = 1–4
Marginally worse
majority-of-5 vs single call, both error directions
−39.8% calls
first-to-3 stopping; identical verdicts to majority-of-5

What this licenses: for this deployed configuration, five-repetition majority voting was not worth the cost. First-to-3 stopping recovered all available savings. What this does not license: unscoped claims about repetition. Provider-side nondeterminism is not excluded by five draws, and temperature-0 is not the same as no variability. Scope is one judge, one corpus, judge stage only.
Result 2 — Exploratory
The attractive metric was not the executable quantity
The natural way to evaluate an alternate judge is to ask: when the primary judge is wrong, how often does the alternate get it right? That quantity — P(alt correct | primary wrong) — looks like operational recovery. It is not.
A runtime policy does not know whether the primary is wrong. It sees two judgments and must decide whether to overturn. The reference-conditioned quantity conditions on a label unavailable at runtime, which is why it runs 20–30 points higher than the precision a deployed policy would actually achieve.
Reference-conditioned
0.543
P(alt correct | primary wrong)
[0.466, 0.618]
Operational (overturn-to-fail)
0.366
15 / 41 overturn-to-fail cases correct
Operational (overturn-to-pass)
0.465
73 / 157 overturn-to-pass cases correct
The gap is 20–30 points. A report that shows only the reference-conditioned figure makes alternate-judge escalation look substantially more useful than a deployed policy would find it. Primary errors split 130 false alarms to 32 missed failures — the pooled recovery metric was dominated by the easier-to-recover stratum.

Predeclaration changed the story
The predeclared evaluator pair (selected before joint outcomes) reached 0.543 pooled recovery. The best post-hoc pair reached 0.759 — in an identically structured report. The predeclared pair ranked last of eight candidates on pooled recovery.
If evaluator selection and metric choice happen after looking at the matrix, it is easy to tell yourself a convincing story. The pair that was locked before outcomes ranked tied first of eight on missed-failure recovery (0.469, 15/32 — four candidates share the value exactly) — the operationally relevant stratum — which was only visible because the choice was already frozen.
Evidence class: pair selection is preregistered (class A); the operational decomposition was conducted after outcomes were visible (exploratory). Both facts are stated.
Result 3 — Held-out validation
Deterministic rules needed their own held-out test
Four deterministic pre-check rules (R1–R4) were preregistered and frozen before evaluation on a 1,260-case held-out arm (1,259 scored; one case had no parseable evaluator verdict). Each fires on production-visible inputs only — no oracle access. Their results were heterogeneous.
| Rule | What it checks | Disposition | What it taught |
|---|---|---|---|
| R1 | Self-contradictory infeasibility claim: the agent calls report_infeasible while asserting completion | EVIDENCE HALF REJECTED | 0.51× — the signal points the opposite way to its specification; 17 of 27 firings are true successes. Had it been given veto power it would have helped 6 and harmed 14. The escalation half survives: 33.3% evaluator error on firings vs 15.2% elsewhere. |
| R2 | Negative self-report on an imperative modification goal | ACCEPTED / UNDERPOWERED | 13 eligible cases, 6 verdicts changed — 6 helped, 0 harmed. Meets its preregistered criterion. Firing precision is 1.06× and the rule-of-three harm bound at n=6 is 39.3%. The acceptance stands; the magnitude is not trustworthy. What it licenses is a larger test, not deployment. |
| R3 | Evidence-gap detection: goal references an image the evaluator cannot observe | PASS | 25.6% evaluator error on 133 fired cases vs 14.4% elsewhere. Clearest positive result — with an important caveat below. |
| R4 | Terminal search-results route | FAIL | 1.20× enrichment for reference-fail is real, but evaluator error is 16.1% on firings vs 15.5% elsewhere — a 0.6-point difference. The rule was measuring agent difficulty, not evaluator reliability. |
R3 in detail: escalation, not correction
R3 fires when the evaluator receives a goal that references an image it cannot observe. The practical interpretation: return UNVERIFIABLE and escalate to human review rather than pretending PASS or FAIL is well-supported. Evaluator error on the 133 fired cases was 25.6% vs 14.4% elsewhere — the rule reliably identifies a high-risk subset.
The caveat: 99 of the 133 escalated cases were already correct. R3 is an error-concentration signal, not a correction mechanism. Its value is a function of human review cost, and it survives because its claim is the weakest of the four: it asserts an evidence gap, not a verdict.
The pattern R3 demonstrates
A rule that asserts a structural gap in the evidence — rather than one that claims to know the verdict — is easier to validate and harder to get subtly wrong. R1 and R4 tried to predict outcomes and failed at it. R1's weaker half — the part that only flags a case for review — is the half that survived. R3 makes only that weaker kind of claim, and passed.
Result 4 — Preregistered / Fresh corpus
A good extractor does not make a bad rule good
RC1 was a static required-conjunct detector: given a task, extract the observable actions an agent must perform, fire if any are missing, and let the evaluator's verdict stand otherwise. Three gates were preregistered before the τ-bench corpus was touched:
- A1: firing volume must be <15% (prevents the detector from dominating evaluation)
- A2-traj: trajectory lift must be ≥1.50×
- A3: label-blind extractor precision must be ≥28/30
A1 — FAIL
25.3%
vs <15% ceiling — 501/1,980 fired
A2-traj — FAIL
1.165×
vs ≥1.50× bar — 1.101× within-task
A3 — PASS
28/30
label-blind extractor precision
The extractor worked. The rule it fed did not. RC1 failed both binding gates on the fresh corpus — not because the detector was parsing incorrectly, but because firing on 25% of trajectories and lifting only 1.165× means the conjunction condition was too permissive to isolate the cases where the evaluator actually needed help.
What the oracle established — and did not

A reconstruction using benchmark-authoritative actions reached 2.339× within-task lift, showing that the latent construct carried signal that RC1 failed to recover. This is not the same as showing that a deployable detector of that signal exists — the oracle uses inputs that a runtime evaluator cannot access. Building one would require new representation work that was not attempted.
A null rule requiring no obligation model at all ("did the agent write anything?") reached 1.171× — above RC1's 1.101×. The null rule was preregistered as disqualified before any computation (volume 24.1% fails A1), so it does not compete with RC1 as a mechanism. It functions as a floor: the production detector should at least clear what requires no semantic machinery.
Why the branch was deliberately closed
RC1 failed both binding gates. The oracle confirms the construct is real. But "the construct has signal" and "a deployable detector can be built" are different claims, separated by representation-design work this branch did not complete. Extending the branch without new representation work would be changing the research question, not answering it.
The research changed its mind
Thirteen process errors were documented across 62 commits. Six were caught before the affected experiment's outcomes were visible; seven were caught afterward — including one that had already appeared in a published report.
Join-key collision (33% of τ-bench records)
A (task_id, trial, domain) key silently collapsed agent identity: 660 of 1,980 τ-bench records were scored against the wrong agent's reward signal. Detected because one qualitative case contained a state the classifier itself could not produce. The report was wrong. The repo records the correction and the corrected values.
Mathematically unpassable gate
A preregistered task-level threshold (1.25×) was computed only after outcomes were visible. When computed, it was mathematically unreachable: 81.8% of tasks held a failing trial, capping the indicator below the threshold by construction. Reclassified as non-decision-bearing rather than treated as a failure gate.
Oracle / null comparison across incompatible outcome variables
An early draft compared oracle lift to null-rule lift using two different outcome variables, inflating one figure by approximately 6×. The same error class (comparing rates across incompatible conditionals) recurred three times, each immediately after the prior instance was corrected. Caught and corrected before publication.
Claim withdrawn from synthesis: “four of five mechanisms failed”
A synthesis-level summary compressed the heterogeneous mechanism dispositions into a ratio that had no enumerated denominator and converted two "weakened / no binary gate" outcomes into failures. Withdrawn during the final internal-consistency audit; the summary now names dispositions mechanically.
Full process-failure audit with all 13 entries, their evidence classes, and whether each was caught before or after outcomes were visible: docs/research_synthesis.md §10.
What survived
Not a framework. A set of distinctions that were non-obvious going in and are hard to recover from offline metrics alone:
- →
Measure the deployed evaluator before designing around hypothetical stochasticity.
Five draws at temperature 0.0 left only 20 of 1,106 cases with any intermediate correctness. The variance to recover was not there. Measurement costs one experiment; a five-copy production system costs ongoing inference spend.
- →
Separate missed failures from false alarms.
Primary errors split 130 false alarms to 32 missed failures. A mechanism that recovers one stratum well may be useless or harmful on the other. Pooled metrics obscure this until you look.
- →
Distinguish reference-conditioned statistics from executable policies.
P(alt correct | primary wrong) is not an operational recovery rate. It conditions on a ground-truth label a runtime policy cannot see. The gap between the two was 20–30 points.
- →
Hold deterministic rules to held-out validation too.
R1 and R4 had inverted or near-zero signal on held-out data. A rule that looked plausible on the design corpus failed when the outcomes were hidden.
- →
Allow UNVERIFIABLE when required evidence is structurally absent.
R3 — the mechanism with the cleanest pass — does not claim a verdict. It claims the evaluator lacked evidence. That is the right claim to make when it is true.
- →
Use cheap baselines before crediting semantic machinery.
A null rule requiring no obligation model reached 1.171× where RC1 reached 1.101×. A semantic detector should clear that floor before being given authority.
- →
Preserve correction history.
Seven of thirteen errors were caught after outcomes were visible. In a study about catching errors in evaluations, the process errors are part of the result.
What I would build differently now
- –Instrument evaluator inputs and provenance before adding more judges — know what information the evaluator is operating on before deciding whether to supplement it.
- –Make missing evidence an explicit state, not a tie-broken verdict.
- –Define false-alarm and missed-failure budgets separately before choosing a mechanism.
- –Evaluate production-deployable decision rules rather than oracle-conditioned recovery statistics.
- –Preregister important evaluation comparisons, including the choice of comparison metric.
- –Make join-key and cardinality assertions part of every benchmark pipeline.
- –Always include a cheap or null baseline before crediting semantic machinery with the lift.
Methodology and audit trail
I started this work expecting to justify a richer evaluator-control architecture. The useful result was narrower: most of the mechanisms I tested did not earn the authority I wanted to give them. The repo preserves the failed gates and corrections because those are part of the result.
Published 2026-09-29 · Four-day research sprint · Two public corpora · No synthetic trajectories · Back to corpus