Jason Stiltner
Research Engineer · First Staff Engineer
Research applied inside a production engineering practice: multi-agent coordination, verifiable behavior, scalable oversight — in service of systems that ship, not the other way around.
First Staff Engineer — AI division of a multi-billion-dollar vertical-SaaS PE firm (24 platforms, six $1B+) · Shipped production AI at HCA Healthcare, the largest US hospital system · Accenture Automation CoE
Focus
Pre-linguistic coordination: how agents cooperate without shared language. Verifiable behavior grounded in observable actions rather than stated intentions.
Verification-centered empirical research: AI-accelerated experimentation, controlled evaluation, production deployment.
Corpus
Grounded Commitment Learning
Multi-agent coordination through verifiable behavioral contracts. Agents commit to observable behaviors rather than inferred mental states—enabling external verification without access to internal representations. Grounded in Hart-Moore incomplete contracts theory (Nobel Prize in Economics, 2016).
40.4% hold-up reduction (95% CI: [37.2%, 43.5%]) · r = -0.972 punishment paradox, p < 0.001
The Archive That Cites Itself
Persistent AI advisors, the people they model, and the limits of “more context”—why a user model should preserve the history of its claims rather than a polished conclusion.
Chat with the Research
Grounded RAG chatbot over this site, red-team hardened. Published eval-gate numbers including the bars still unmet, and the real defects the harness found and root-caused—rate limiter, citation injection, retrieval crowding.
Collaborative Nested Learning
Extension of Google Research’s nested optimization: 5 timescales with 9 bidirectional knowledge bridges. Addresses catastrophic interference where fast learning degrades slow-learned representations, via normalization constraints that preserve component distinctiveness during optimization.
+89% accuracy at high regularization, where baseline collapses
Aegis
Systems-architecture layer beneath agent frameworks: durability, verification, and policy—not orchestration. Event-sourced state for resume/replay from any checkpoint, a tool gateway enforcing policy at invocation time, and GCL commitments as first-class objects with explicit failure modes.
303 tests passing
Methods
- Empirical research. AI-accelerated hypothesis generation and experimental iteration, gated by controlled evaluation, statistical validation, and reproducible evidence. Model-generated explanations are hypotheses, not evidence.
- Mechanistic validation. Ablations and interventions that separate predictive success from causal explanation. The punishment paradox is re-derived by CI on every push; CNL's bridge ablation weekly, at full scale — both with the run behind them.
- Formal foundations. Mathematical proofs where applicable. Convergence guarantees, conservation laws, contraction mappings.
- Executable research. Research artifacts built as software: automated evaluation, reproducible experiments, CI, test coverage, inspectable results.
- Production systems. Research shaped by the constraints of systems that ship. Observability, failure recovery, deployment — GCP, Terraform, Docker, HIPAA-compliant architectures, multi-provider routing, edge inference.
Background
Path here: language, then automation, then ML.
M.A. Université de Paris VII (French-language graduate program)
Littérature, Langues, et Civilisations des Pays Anglophones