Jason Stiltner

Research Engineer · Staff Engineer at GSV AI Labs

The role is staff engineer; the work is at the boundary of production AI systems and empirical evaluation — agent infrastructure, the behavioral evaluation systems that decide whether it works, and controlled experiments on whether those evaluators can be trusted to decide.

Staff Engineer at GSV AI Labs, the AI division of private equity firm Greater Sum Ventures — building CharlieIQ · Shipped production AI at HCA Healthcare, the largest US hospital system · Accenture Automation CoE

jason@jasonstiltner.com · GitHub · LinkedIn

Focus

Verification-centered empirical research: AI-accelerated experimentation, controlled evaluation, production deployment.

The work runs as a loop: production behavior exposes what the evals missed, an evaluation system is built to catch it, and the evaluator itself is then tested against held-out and fresh data before it is trusted to decide anything. Nulls, falsifications and corrections stay published — they are what makes the rest checkable. Feeding the result back into architecture is the step still argued for rather than built.

Pre-linguistic coordination: how agents cooperate without shared language. Verifiable behavior grounded in observable actions rather than stated intentions.

Corpus

Evaluator Assurance: What Survived

When can a model-based evaluator be trusted to decide whether an agent succeeded — to retry, escalate, or release? Falsification study testing seven evaluator-assurance mechanisms on two real agentic-trajectory corpora. Repeated execution, alternate-judge pairing, four deterministic rules, and a static required-conjunct matcher were preregistered and tested against held-out and fresh-corpus data. Most failed or weakened; the clearest positive result was an evidence-gap escalation signal.

25.6% vs 14.4% evaluator error on R3-fired vs unfired cases (held-out, n=1,260) · 0.366 / 0.465 operational overturn precision vs 0.543 reference-conditioned · −39.8% calls first-to-3 stopping with zero verdict change vs majority-of-5

Branch closed 2026-09-29. Thirteen process errors documented, six caught before the affected experiment's outcomes were visible. All corrections are recorded in the repo. The static required-conjunct rule (RC1) failed both binding gates on fresh τ-bench data; an oracle reconstruction showed the latent construct carried signal (2.339×) that the production detector failed to recover.

When Production Disagrees with the Architecture

A caller answered two questions when our voice agent had asked one. That small production failure exposed a larger problem: evals can preserve counterexamples without preserving what they should teach us about architecture. Production evals → architectural assumptions → prospective evidence → earned autonomy.

The evaluation machinery described in Act I is deployed. The architectural-learning loop the essay argues for is a proposal — nothing in Act III has been built.

Grounded Commitment Learning

Multi-agent coordination through verifiable behavioral contracts. Agents commit to observable behaviors rather than inferred mental states, so a third party can check compliance without access to internal representations. The framework is formalized against Hart-Moore incomplete contracts theory; nearly all of the evidence below is simulation, and the two figures from real LLM agents are listed first.

4/4 agents better calibrated by external assessment than self-report (real LLMs) · +0.017 motivation effect in real LLMs — a null, p = 0.56, CI excludes the predicted +0.065 · 40.4% hold-up reduction in simulation (95% CI: [37.2%, 43.5%])

Corrected in place repeatedly, most recently September 2026: several published figures were traced to hardcoded chart constants rather than experiment output, and the self-selection result was falsified by an adversarial re-audit, re-measured, and then failed to reproduce in real LLM agents. The original wording and the defect are kept on the page. Only thirteen of the experiment scripts import the GCL package; everything from 19 onward is a standalone simulation.

Chat with the Research

Grounded RAG chatbot over this site, red-team hardened. Published eval-gate numbers including the bars still unmet, and the real defects the harness found and root-caused—rate limiter, citation injection, retrieval crowding.

The Bottleneck Moves

Coding agents made implementation cheap, so the bottleneck was supposed to move upstream into specification and architecture. Some of it did; more of it accumulated downstream in convergence — restoring mergeability, asynchronous review round trips, serialized resources, and acceptance criteria the executor cannot verify. Drawn from a retrospective of twenty merged changes in a production system, none of which a reader can check.

The only entry here with no checkable evidence. The underlying sample is employer work product: the repository is not public and the methodology note is not mine to publish, so every figure in the essay rests on my word alone. Stated at the top of the page rather than at the bottom.

Full corpus — 17 entries, filterable by facet →

Methods

  1. Empirical research. AI-accelerated hypothesis generation and experimental iteration, gated by controlled evaluation, statistical validation, and reproducible evidence. Model-generated explanations are hypotheses, not evidence.
  2. Mechanistic validation. Ablations and interventions that separate predictive success from causal explanation. The punishment paradox is re-derived by CI on every push; CNL's bridge ablation weekly, at full scale — both with the run behind them.
  3. Formal foundations. Mathematical proofs where applicable. Convergence guarantees, conservation laws, contraction mappings.
  4. Executable research. Research artifacts built as software: automated evaluation, reproducible experiments, CI, test coverage, inspectable results.
  5. Production systems. Research shaped by the constraints of systems that ship. Observability, failure recovery, deployment — GCP, Terraform, Docker, HIPAA-compliant architectures, multi-provider routing, edge inference.

Background

Path here: language, then automation, then ML.

M.A. Université de Paris VII (French-language graduate program)
Littérature, Langues, et Civilisations des Pays Anglophones

More →