<!-- generated by scripts/generate-agent-docs.ts -->

> **Staff AI Engineer — évaluation d'agents et fiabilité des évaluateurs**
> Je construis les systèmes qui permettent de savoir si une IA en production fonctionne réellement : les évaluations qui détectent les défaillances, l'outillage qui les explique, et les tests qui vérifient si les évaluateurs eux-mêmes méritent qu'on leur fasse confiance. Basé à Nashville.
>
> Source: https://jasonstiltner.com/fr/

---

# Jason Stiltner

Staff AI Engineer  Évaluation d'agents et fiabilité des évaluateurs

Je construis les systèmes qui permettent de savoir si une IA en production fonctionne réellement : les évaluations qui détectent les défaillances, l'outillage qui les explique, et les tests qui vérifient si les évaluateurs eux-mêmes méritent qu'on leur fasse confiance.

Staff Engineer chez CharlieIQ, le laboratoire d'IA de Greater Sum Ventures · IA de production déployée chez HCA Healthcare, le plus grand groupe hospitalier américain · Accenture, Automation Center of Excellence

[jason@jasonstiltner.com](mailto:jason@jasonstiltner.com) · [GitHub](https://github.com/jstiltner) · [LinkedIn](https://linkedin.com/in/jasonlstiltner)

## Approche

Y a-t-il eu défaillance ?

Des évaluations comportementales qui détectent ce que les tests unitaires laissent passer — [When Production Disagrees](https://jasonstiltner.com/writing/when-production-disagrees/) et les [campagnes d'évaluation publiées](https://jasonstiltner.com/tools/research-chat/evals/) du chatbot de ce site.

Pourquoi a-t-elle eu lieu ?

Un outillage qui transforme une trajectoire en échec en explication — [le prototype Stage 03](https://jasonstiltner.com/writing/when-production-disagrees/#stage-03-spike), dont le premier résultat ne s'est pas reproduit.

Peut-on faire confiance au juge ?

Des tests de falsification sur les évaluateurs eux-mêmes — [Evaluator Assurance](https://jasonstiltner.com/projects/evaluator-assurance/), sept mécanismes sur deux corpus d'agents publics, dont aucun n'a survécu.

Les résultats nuls, les falsifications et les corrections restent publiés — c'est ce qui rend le reste vérifiable.

## Corpus

Les cinq premières entrées du corpus complet, dans le même ordre. Les résumés restent en anglais — ce sont les mêmes extraits que ceux des pages projet, pour ne rien reformuler en les traduisant.

### [Evaluator Assurance: What Survived](https://jasonstiltner.com/projects/evaluator-assurance/)

Falsification study of seven evaluator-assurance mechanisms on two public agent corpora. None survived as a robust improvement: repetition had nothing to average over, alternate-judge pairing's headline statistic turned out to be the wrong conditional, and the one rule that passed its numeric gate cleanly — the R3 evidence-gap rule — shrank from a +11.2 pp separation to +5.8 pp (p = 0.26) once stratified by environment.

1.8% of cases had five draws that were not unanimous — 20 of 1,106 cases, 5 calls each, one judge at temperature 0, judge stage only · 0.366 / 0.465 operational overturn precision, beside 0.543 recovery pooled over both error directions · +11.2 pp → +5.8 pp the R3 evidence-gap rule's held-out separation, before and after conditioning on environment (p = 0.26)

Branch closed 2026-09-29, then corrected in place after a separate end-to-end reproduction found eleven wrong or over-strong figures — including the one this entry used to lead with. The full record, with the original wording kept beside each correction, is on the project page.

### [When Production Disagrees with the Architecture](https://jasonstiltner.com/writing/when-production-disagrees/)

A caller answered two questions when our voice agent had asked one. That small production failure exposed a larger problem: evals can preserve counterexamples without preserving what they should teach us about architecture. Production evals → architectural assumptions → prospective evidence → earned autonomy.

The evaluation machinery described in Act I is deployed. The architectural-learning loop the essay argues for is a proposal — nothing in Act III has been built.

### [Explaining failures](https://jasonstiltner.com/writing/when-production-disagrees/#stage-03-spike)

A bounded implementation of the Explain stage, run against preserved failures from this site's evaluation corpus: it generates competing explanations, requests evidence through a typed allowlist of read-only inspections, and revises its hypotheses from what it observes. The first calibration result did not survive being frozen and rerun.

4/6 → 2/30 calibration result before and after the system was frozen and rerun — 28 of 30 trials remained unresolved, with no incorrect resolution in this small sample · 0/5 held-out failures resolved, in part because the two available probes could not reach most of the evidence the system wanted to inspect

A spike, not a system: one real case moved from unresolved to a named explanation after deterministic inspection identified a grading-rule change, and one became narrower and stayed unresolved. Nothing here is deployed or in a decision path.

### [Chat with the Research](https://jasonstiltner.com/projects/chat-with-the-research/)

Grounded RAG chatbot over this site, red-team hardened. Published eval-gate numbers including the bars still unmet, and the real defects the harness found and root-caused—rate limiter, citation injection, retrieval crowding.

### [Grounded Commitment Learning](https://jasonstiltner.com/projects/grounded-commitment-learning/)

Multi-agent coordination through verifiable behavioral contracts. Agents commit to observable behaviors rather than inferred mental states, so a third party can check compliance without access to internal representations. The framework is formalized against Hart-Moore incomplete contracts theory; nearly all of the evidence below is simulation, and the two figures from real LLM agents are listed first.

4/4 agents better calibrated by external assessment than self-report (real LLMs) · +0.017 motivation effect in real LLMs — a null, p = 0.56, CI excludes the predicted +0.065 · 40.4% hold-up reduction in simulation (95% CI: \[37.2%, 43.5%\])

Corrected in place repeatedly, most recently September 2026: several published figures were traced to hardcoded chart constants rather than experiment output, and the self-selection result was falsified by an adversarial re-audit, re-measured, and then failed to reproduce in real LLM agents. The original wording and the defect are kept on the page. Only thirteen of the experiment scripts import the GCL package; everything from 19 onward is a standalone simulation.

[Corpus complet — 16 entrées, filtrables par facette →](https://jasonstiltner.com/corpus/)

## Méthode

1.  Recherche empirique. L'IA accélère la génération d'hypothèses et l'itération expérimentale, sous condition d'une évaluation contrôlée, d'une validation statistique et de preuves reproductibles. Une explication produite par un modèle reste une hypothèse, pas une preuve.
2.  Validation mécanistique. Des ablations et interventions qui distinguent un succès prédictif d'une explication causale. Le paradoxe de la punition est re-dérivé par la CI à chaque push ; l'ablation des ponts de CNL, chaque semaine, à pleine échelle — [preuves à l'appui pour les deux](https://jasonstiltner.com/corpus/reproducibility/).
3.  Fondements formels. Des preuves mathématiques quand elles s'appliquent : garanties de convergence, lois de conservation, applications contractantes.
4.  Recherche exécutable. Les artefacts de recherche construits comme du logiciel : évaluation automatisée, expériences reproductibles, CI, couverture de tests, résultats inspectables.
5.  Systèmes en production. Une recherche façonnée par les contraintes des systèmes qui tiennent en production : observabilité, reprise après incident, déploiement — GCP, Terraform, Docker, architectures conformes HIPAA, routage multi-fournisseurs, inférence en edge.

## Parcours

Le chemin jusqu'ici : d'abord la langue, puis l'automatisation, puis le machine learning.

Master, Université Paris Cité (anciennement Paris VII) — cursus francophone
Littérature, Langues, et Civilisations des Pays Anglophones

[En savoir plus →](https://jasonstiltner.com/fr/about/)
