Jason Stiltner

Research Engineer · Staff Engineer at GSV AI Labs

The role is staff engineer; the work is at the boundary of production AI systems and empirical evaluation — agent infrastructure, the behavioral evaluation systems that decide whether it works, and controlled experiments on whether those evaluators can be trusted to decide.

Staff Engineer at GSV AI Labs, the AI division of private equity firm Greater Sum Ventures — building CharlieIQ · Shipped production AI at HCA Healthcare, the largest US hospital system · Accenture Automation CoE

jason@jasonstiltner.com · GitHub · LinkedIn

Focus

Verification-centered empirical research: AI-accelerated experimentation, controlled evaluation, production deployment.

The work runs as a loop: production behavior exposes what the evals missed, an evaluation system is built to catch it, and the evaluator itself is then tested against held-out and fresh data before it is trusted to decide anything. Nulls, falsifications and corrections stay published — they are what makes the rest checkable. Feeding the result back into architecture is the step still argued for rather than built.

Pre-linguistic coordination: how agents cooperate without shared language. Verifiable behavior grounded in observable actions rather than stated intentions.

Corpus

Evaluator Assurance: What Survived

Falsification study of seven evaluator-assurance mechanisms on two public agent corpora. None survived as a robust improvement: repetition had nothing to average over, alternate-judge pairing's headline statistic turned out to be the wrong conditional, and the one rule that passed its numeric gate cleanly — the R3 evidence-gap rule — shrank from a +11.2 pp separation to +5.8 pp (p = 0.26) once stratified by environment.

1.8% of cases had five draws that were not unanimous — 20 of 1,106 cases, 5 calls each, one judge at temperature 0, judge stage only · 0.366 / 0.465 operational overturn precision, beside 0.543 recovery pooled over both error directions · +11.2 pp → +5.8 pp the R3 evidence-gap rule's held-out separation, before and after conditioning on environment (p = 0.26)

Branch closed 2026-09-29. Corrected 2026-10-02 after an independent end-to-end reproduction: this entry summarised R3, the evidence-gap escalation rule, as the result that survived — superseded, because all 133 of R3's firings fall in the single benchmark carrying the highest judge-error point estimate (22.4%, next-highest 19.8%), so the preregistered test compared one slice against three others. Conditioned on environment the separation is +5.8 pp, p = 0.26, against +11.2 pp, p = 0.0015 pooled; R3's preregistered disposition is unchanged and the residual stays positive, so what shrinks is the size of the claim — narrowed, not withdrawn. The same reproduction found ten further wrong or over-strong figures and protocol descriptions, including a pooled-vs-direction-specific subtraction behind the published metric gap. Two process-error counts, kept apart: 13 found during the study (6 before the relevant outcomes were visible, 7 afterwards) and 11 more found by the external reproduction, which no safeguard here caught. All corrections are recorded in the repo. The static required-conjunct rule (RC1) failed both binding gates on fresh τ-bench data; an oracle reconstruction showed the latent construct carried signal (2.339×) that the production detector failed to recover.

When Production Disagrees with the Architecture

A caller answered two questions when our voice agent had asked one. That small production failure exposed a larger problem: evals can preserve counterexamples without preserving what they should teach us about architecture. Production evals → architectural assumptions → prospective evidence → earned autonomy.

The evaluation machinery described in Act I is deployed. The architectural-learning loop the essay argues for is a proposal — nothing in Act III has been built.

Grounded Commitment Learning

Multi-agent coordination through verifiable behavioral contracts. Agents commit to observable behaviors rather than inferred mental states, so a third party can check compliance without access to internal representations. The framework is formalized against Hart-Moore incomplete contracts theory; nearly all of the evidence below is simulation, and the two figures from real LLM agents are listed first.

4/4 agents better calibrated by external assessment than self-report (real LLMs) · +0.017 motivation effect in real LLMs — a null, p = 0.56, CI excludes the predicted +0.065 · 40.4% hold-up reduction in simulation (95% CI: [37.2%, 43.5%])

Corrected in place repeatedly, most recently September 2026: several published figures were traced to hardcoded chart constants rather than experiment output, and the self-selection result was falsified by an adversarial re-audit, re-measured, and then failed to reproduce in real LLM agents. The original wording and the defect are kept on the page. Only thirteen of the experiment scripts import the GCL package; everything from 19 onward is a standalone simulation.

Chat with the Research

Grounded RAG chatbot over this site, red-team hardened. Published eval-gate numbers including the bars still unmet, and the real defects the harness found and root-caused—rate limiter, citation injection, retrieval crowding.

The Bottleneck Moves

Coding agents made implementation cheap, so the bottleneck was supposed to move upstream into specification and architecture. Some of it did; more of it accumulated downstream in convergence — restoring mergeability, asynchronous review round trips, serialized resources, and acceptance criteria the executor cannot verify. Drawn from a retrospective of twenty merged changes in a production system, none of which a reader can check.

The only entry here with no checkable evidence. The underlying sample is employer work product: the repository is not public and the methodology note is not mine to publish, so every figure in the essay rests on my word alone. Stated at the top of the page rather than at the bottom.

Full corpus — 17 entries, filterable by facet →

Methods

  1. Empirical research. AI-accelerated hypothesis generation and experimental iteration, gated by controlled evaluation, statistical validation, and reproducible evidence. Model-generated explanations are hypotheses, not evidence.
  2. Mechanistic validation. Ablations and interventions that separate predictive success from causal explanation. The punishment paradox is re-derived by CI on every push; CNL's bridge ablation weekly, at full scale — both with the run behind them.
  3. Formal foundations. Mathematical proofs where applicable. Convergence guarantees, conservation laws, contraction mappings.
  4. Executable research. Research artifacts built as software: automated evaluation, reproducible experiments, CI, test coverage, inspectable results.
  5. Production systems. Research shaped by the constraints of systems that ship. Observability, failure recovery, deployment — GCP, Terraform, Docker, HIPAA-compliant architectures, multi-provider routing, edge inference.

Background

Path here: language, then automation, then ML.

M.A. Université de Paris VII (French-language graduate program)
Littérature, Langues, et Civilisations des Pays Anglophones

More →