<!-- generated by scripts/generate-agent-docs.ts -->

> **Evaluator Assurance: What Survived**
> Falsification study testing seven evaluator-assurance mechanisms on two real agentic-trajectory corpora (AgentRewardBench, n=1,106; τ-bench, n=1,980), with preregistration, held-out validation, fresh-corpus testing, and a full process-failure audit. Most mechanisms weakened or failed; the clearest positive result was an evidence-gap escalation signal (R3: 25.6% vs 14.4% evaluator error). The static required-conjunct rule (RC1) failed both binding gates on fresh data despite a 28/30 extractor audit.
>
> Source: https://jasonstiltner.com/projects/evaluator-assurance/

---

*Part of my research on [robust evaluation of adaptive systems](https://jasonstiltner.com/corpus/). Branch closed 2026-09-29.*

AI Evaluation · Agent Reliability · Experimental Systems

# Evaluator Assurance: What Survived

Published Sep 29, 2026

I tested seven plausible ways to make evaluator-driven AI systems more trustworthy across AgentRewardBench and τ-bench. Most weakened or failed when exposed to preregistration, held-out data, operational metrics, and cheap baselines. The useful result was narrower than the architecture I started with.

This is a falsification study, not a framework proposal. The research value comes from determining what did not earn authority, and preserving enough evidence that a skeptical reader can inspect why.

Production-deployed— two public corpora — no synthetic trajectories

3,086

trajectory records across two studies

5,530

paid repeat judgments (R=5)

$46.18

measured inference spend

7

mechanisms frozen and tested

13

documented process errors

4-day

intensive research sprint

[View repo](https://github.com/jstiltner/assurance-planner) · [Jump to findings](#findings)

## The question

LLM-backed systems increasingly use evaluators to decide whether an agent succeeded, whether to retry, whether to escalate, or whether to trust a generated result. When those evaluators make mistakes, the intuitive response is often: sample more, add another judge, add routing, add deterministic guards, combine everything.

The project asks a more precise question:

> Which of those mechanisms actually earn their cost and authority when tested against real trajectories, with preregistered gates and a held-out corpus they have never seen?

### Corpora

| Corpus | n | Evidence class |
| --- | --- | --- |
| AgentRewardBench
rev b6d17e6 | 1,106 | Preregistered |
| τ-bench
historical\_trajectories/ | 1,980 | Fresh / never touched |

### Mechanisms tested

-   M1Repeated execution of the same judge
-   M2Alternate-evaluator pairing / escalation
-   M3–M6Deterministic pre-check rules (R1–R4, held-out)
-   M7Static required-conjunct matching (RC1, fresh corpus)

Result 1 — Preregistered

## Repetition had almost nothing to average over

gpt-4o-2024-11-20 · temperature 0.0 · seed 0 · judge stage only

Five repeated calls were made on each of 1,106 AgentRewardBench trajectories — 5,530 calls total. If repeated sampling were recovering meaningful variance, a substantial fraction of cases would land in the middle of the k=0–5 distribution.

![Bar chart showing k-distribution (0–5 correct out of 5 repetitions) for reference-fail and reference-success cases. Both strata are sharply bimodal: the vast majority of cases land at k=0 or k=5, with only 20 of 1,106 cases falling in k=1–4.](https://jasonstiltner.com/projects/evaluator-assurance/fig1_k_distribution.png)

**Figure 1.** k-distribution across R=5 repetitions. Of 1,106 cases, only 20 had intermediate correctness (k=1–4); 98.8% of reference-fail cases and 96.6% of reference-success cases landed at k=0 or k=5.

The deployed judge was not perfectly deterministic. It was simply stable enough that repeated sampling had almost no useful error to recover. Only 20 of 1,106 cases showed any intermediate correctness across five draws.

20

of 1,106 cases had k = 1–4

Marginally worse

majority-of-5 vs single call, both error directions

−39.8% calls

first-to-3 stopping; identical verdicts to majority-of-5

![Bar chart comparing evaluator error rates across voting policies: single call, majority-of-3, majority-of-5, and first-to-3 stopping. Majority voting is marginally worse than a single call. First-to-3 stopping matches majority-of-5 exactly.](https://jasonstiltner.com/projects/evaluator-assurance/fig2_policy_error_comparison.png)

**Figure 2.** Error rate by voting policy. First-to-3 stopping reproduced majority-of-5 verdicts exactly at 39.8% fewer calls ($27.86 vs $46.18).

**What this licenses:** for this deployed configuration, five-repetition majority voting was not worth the cost. First-to-3 stopping recovered all available savings. **What this does not license:** unscoped claims about repetition. Provider-side nondeterminism is not excluded by five draws, and temperature-0 is not the same as no variability. Scope is one judge, one corpus, judge stage only.

Result 2 — Exploratory

## The attractive metric was not the executable quantity

The natural way to evaluate an alternate judge is to ask: when the primary judge is wrong, how often does the alternate get it right? That quantity — P(alt correct | primary wrong) — looks like operational recovery. It is not.

A runtime policy does not know whether the primary is wrong. It sees two judgments and must decide whether to overturn. The reference-conditioned quantity conditions on a label unavailable at runtime, which is why it runs 20–30 points higher than the precision a deployed policy would actually achieve.

Reference-conditioned

0.543

P(alt correct | primary wrong)
\[0.466, 0.618\]

Operational (overturn-to-fail)

0.366

15 / 41 overturn-to-fail cases correct

Operational (overturn-to-pass)

0.465

73 / 157 overturn-to-pass cases correct

The gap is 20–30 points. A report that shows only the reference-conditioned figure makes alternate-judge escalation look substantially more useful than a deployed policy would find it. Primary errors split 130 false alarms to 32 missed failures — the pooled recovery metric was dominated by the easier-to-recover stratum.

![Bar chart comparing reference-conditioned recovery (0.543) against operational overturn-to-fail precision (0.366) and overturn-to-pass precision (0.465). A bracket labels the 20–30 point gap.](https://jasonstiltner.com/projects/evaluator-assurance/fig4_ref_vs_operational.png)

**Figure 4.** Reference-conditioned recovery vs operational overturn precision. The reference-conditioned quantity conditions on a ground-truth label; an operational policy must decide without it.

### Predeclaration changed the story

The predeclared evaluator pair (selected before joint outcomes) reached 0.543 pooled recovery. The best post-hoc pair reached 0.759 — in an identically structured report. The predeclared pair ranked last of eight candidates on pooled recovery.

If evaluator selection and metric choice happen after looking at the matrix, it is easy to tell yourself a convincing story. The pair that was locked before outcomes ranked *tied first* of eight on missed-failure recovery (0.469, 15/32 — four candidates share the value exactly) — the operationally relevant stratum — which was only visible because the choice was already frozen.

**Evidence class:** pair selection is preregistered (class A); the operational decomposition was conducted after outcomes were visible (exploratory). Both facts are stated.

Result 3 — Held-out validation

## Deterministic rules needed their own held-out test

Four deterministic pre-check rules (R1–R4) were preregistered and frozen before evaluation on a 1,260-case held-out arm (1,259 scored; one case had no parseable evaluator verdict). Each fires on production-visible inputs only — no oracle access. Their results were heterogeneous.

| Rule | What it checks | Disposition | What it taught |
| --- | --- | --- | --- |
| R1 | Self-contradictory infeasibility claim: the agent calls report\_infeasible while asserting completion | EVIDENCE HALF REJECTED | 0.51× — the signal points the opposite way to its specification; 17 of 27 firings are true successes. Had it been given veto power it would have helped 6 and harmed 14. The escalation half survives: 33.3% evaluator error on firings vs 15.2% elsewhere. |
| R2 | Negative self-report on an imperative modification goal | ACCEPTED / UNDERPOWERED | 13 eligible cases, 6 verdicts changed — 6 helped, 0 harmed. Meets its preregistered criterion. Firing precision is 1.06× and the rule-of-three harm bound at n=6 is 39.3%. The acceptance stands; the magnitude is not trustworthy. What it licenses is a larger test, not deployment. |
| R3 | Evidence-gap detection: goal references an image the evaluator cannot observe | PASS | 25.6% evaluator error on 133 fired cases vs 14.4% elsewhere. Clearest positive result — with an important caveat below. |
| R4 | Terminal search-results route | FAIL | 1.20× enrichment for reference-fail is real, but evaluator error is 16.1% on firings vs 15.5% elsewhere — a 0.6-point difference. The rule was measuring agent difficulty, not evaluator reliability. |

### R3 in detail: escalation, not correction

R3 fires when the evaluator receives a goal that references an image it cannot observe. The practical interpretation: return UNVERIFIABLE and escalate to human review rather than pretending PASS or FAIL is well-supported. Evaluator error on the 133 fired cases was 25.6% vs 14.4% elsewhere — the rule reliably identifies a high-risk subset.

The caveat: 99 of the 133 escalated cases were already correct. R3 is an error-concentration signal, not a correction mechanism. Its value is a function of human review cost, and it survives because its claim is the weakest of the four: it asserts an evidence gap, not a verdict.

The pattern R3 demonstrates

A rule that asserts a structural gap in the evidence — rather than one that claims to know the verdict — is easier to validate and harder to get subtly wrong. R1 and R4 tried to predict outcomes and failed at it. R1's weaker half — the part that only flags a case for review — is the half that survived. R3 makes only that weaker kind of claim, and passed.

Result 4 — Preregistered / Fresh corpus

## A good extractor does not make a bad rule good

RC1 was a static required-conjunct detector: given a task, extract the observable actions an agent must perform, fire if any are missing, and let the evaluator's verdict stand otherwise. Three gates were preregistered before the τ-bench corpus was touched:

-   A1: firing volume must be <15% (prevents the detector from dominating evaluation)
-   A2-traj: trajectory lift must be ≥1.50×
-   A3: label-blind extractor precision must be ≥28/30

A1 — FAIL

25.3%

vs <15% ceiling — 501/1,980 fired

A2-traj — FAIL

1.165×

vs ≥1.50× bar — 1.101× within-task

A3 — PASS

28/30

label-blind extractor precision

The extractor worked. The rule it fed did not. RC1 failed both binding gates on the fresh corpus — not because the detector was parsing incorrectly, but because firing on 25% of trajectories and lifting only 1.165× means the conjunction condition was too permissive to isolate the cases where the evaluator actually needed help.

### What the oracle established — and did not

![Bar chart comparing within-task lift for three conditions: oracle reconstruction (2.339×, benchmark-authoritative), null rule (1.171×, no obligation model), and RC1 production detector (1.101×). Labels identify evidence class.](https://jasonstiltner.com/projects/evaluator-assurance/fig3_lift_comparison.png)

**Figure 3.** Within-task lift comparison. Oracle: 2.339× (benchmark-authoritative actions, undeployable). Null rule: 1.171× (no obligation model, preregistered as disqualified). RC1: 1.101× (production-visible, fresh corpus).

A reconstruction using benchmark-authoritative actions reached 2.339× within-task lift, showing that the latent construct carried signal that RC1 failed to recover. This is not the same as showing that a deployable detector of that signal exists — the oracle uses inputs that a runtime evaluator cannot access. Building one would require new representation work that was not attempted.

A null rule requiring no obligation model at all ("did the agent write anything?") reached 1.171× — above RC1's 1.101×. The null rule was preregistered as disqualified before any computation (volume 24.1% fails A1), so it does not compete with RC1 as a mechanism. It functions as a floor: the production detector should at least clear what requires no semantic machinery.

Why the branch was deliberately closed

RC1 failed both binding gates. The oracle confirms the construct is real. But "the construct has signal" and "a deployable detector can be built" are different claims, separated by representation-design work this branch did not complete. Extending the branch without new representation work would be changing the research question, not answering it.

## The research changed its mind

Thirteen process errors were documented across 62 commits. Six were caught before the affected experiment's outcomes were visible; seven were caught afterward — including one that had already appeared in a published report.

### Join-key collision (33% of τ-bench records)

A (task\_id, trial, domain) key silently collapsed agent identity: 660 of 1,980 τ-bench records were scored against the wrong agent's reward signal. Detected because one qualitative case contained a state the classifier itself could not produce. The report was wrong. The repo records the correction and the corrected values.

### Mathematically unpassable gate

A preregistered task-level threshold (1.25×) was computed only after outcomes were visible. When computed, it was mathematically unreachable: 81.8% of tasks held a failing trial, capping the indicator below the threshold by construction. Reclassified as non-decision-bearing rather than treated as a failure gate.

### Oracle / null comparison across incompatible outcome variables

An early draft compared oracle lift to null-rule lift using two different outcome variables, inflating one figure by approximately 6×. The same error class (comparing rates across incompatible conditionals) recurred three times, each immediately after the prior instance was corrected. Caught and corrected before publication.

### Claim withdrawn from synthesis: “four of five mechanisms failed”

A synthesis-level summary compressed the heterogeneous mechanism dispositions into a ratio that had no enumerated denominator and converted two "weakened / no binary gate" outcomes into failures. Withdrawn during the final internal-consistency audit; the summary now names dispositions mechanically.

Full process-failure audit with all 13 entries, their evidence classes, and whether each was caught before or after outcomes were visible: [docs/research\_synthesis.md §10](https://github.com/jstiltner/assurance-planner/blob/main/docs/research_synthesis.md).

## What survived

Not a framework. A set of distinctions that were non-obvious going in and are hard to recover from offline metrics alone:

-   →

    Measure the deployed evaluator before designing around hypothetical stochasticity.

    Five draws at temperature 0.0 left only 20 of 1,106 cases with any intermediate correctness. The variance to recover was not there. Measurement costs one experiment; a five-copy production system costs ongoing inference spend.

-   →

    Separate missed failures from false alarms.

    Primary errors split 130 false alarms to 32 missed failures. A mechanism that recovers one stratum well may be useless or harmful on the other. Pooled metrics obscure this until you look.

-   →

    Distinguish reference-conditioned statistics from executable policies.

    P(alt correct | primary wrong) is not an operational recovery rate. It conditions on a ground-truth label a runtime policy cannot see. The gap between the two was 20–30 points.

-   →

    Hold deterministic rules to held-out validation too.

    R1 and R4 had inverted or near-zero signal on held-out data. A rule that looked plausible on the design corpus failed when the outcomes were hidden.

-   →

    Allow UNVERIFIABLE when required evidence is structurally absent.

    R3 — the mechanism with the cleanest pass — does not claim a verdict. It claims the evaluator lacked evidence. That is the right claim to make when it is true.

-   →

    Use cheap baselines before crediting semantic machinery.

    A null rule requiring no obligation model reached 1.171× where RC1 reached 1.101×. A semantic detector should clear that floor before being given authority.

-   →

    Preserve correction history.

    Seven of thirteen errors were caught after outcomes were visible. In a study about catching errors in evaluations, the process errors are part of the result.

## What I would build differently now

-   Instrument evaluator inputs and provenance before adding more judges — know what information the evaluator is operating on before deciding whether to supplement it.
-   Make missing evidence an explicit state, not a tie-broken verdict.
-   Define false-alarm and missed-failure budgets separately before choosing a mechanism.
-   Evaluate production-deployable decision rules rather than oracle-conditioned recovery statistics.
-   Preregister important evaluation comparisons, including the choice of comparison metric.
-   Make join-key and cardinality assertions part of every benchmark pipeline.
-   Always include a cheap or null baseline before crediting semantic machinery with the lift.
