Part of my research on robust evaluation of adaptive systems.
Grounded Commitment Learning
Coordination Without Shared Semantics
Updated Sep 21, 2026Natural language coordination assumes shared semantics — that "complete the task" means the same thing to all agents. That assumption fails under semantic drift: agents with different training, different architectures, or the same agent at two points in time can read identical phrases differently. Rather than trying to solve representation, GCL asks whether coordination can be grounded in externally observable consequences instead: a commitment's meaning is fixed not by how agents interpret it, but by which observable outcomes count as success and which count as failure.
The motivating question is a scalable-oversight one — how do you verify agent behaviour when you cannot inspect internal states? Behavioural verifiability inside a designed contract framework is what this line of work is aimed at. It is a research direction with a partial simulation record behind it, not a solved problem, and the sections below are organised so you can tell which is which.
The message-count win comes with a loss on the same experiment's primary metric: GCL places last of six on coordination efficiency (0.645 vs CNP's 0.824) and is not on the Pareto frontier. Two further badges — sample efficiency and coordination quality vs. MARL — were withdrawn in September 2026. The research record lists every correction, and what currently survives separates the claims by how much evidence is behind them.
Everything on this page except Experiments 41 and 41b is simulation. Reproduce it: github.com/jstiltner/gcl.
Research record
This page has been corrected in place 11 times since August 2026, most recently 2026-09-21. Nothing below has been removed — each link goes to the finding it replaces. For what the surviving claims are, rather than what was withdrawn, see what currently survives.
- RetractedSelf-selection beats external assignment by 81%, driven by information asymmetry· Aug 2026
- RetractedFour network-property figures (82.3% convergence, 0.699 clustering, 0.745 Gini, 26.5% efficiency)· Sep 2026
- SupersededRedemption effect size (39.3% → 60.0%, +52.7%)· Sep 2026
- RetractedCoordination-scaling claims (Dunbar-like limit, message growth, network topology)· Sep 2026
- Qualified“25–50× more sample-efficient, 97% of MARL quality” vs. MARL· Sep 2026
- RetractedMAS comparison chart coordinates (competitor message counts and success rates)· Sep 2026
- RetractedTemplate-sharing figures (+17% cooperation, Gini 0.35 → 0.25, bottom quartile 34% faster, random +8%)· Sep 2026
- QualifiedGCL as a “foundation” for scalable oversight and constitutional-AI-style guardrails· Sep 2026
- RetractedGaming rates by reputation condition (12% / 45% / 18% / 15%) and the statistics t = 8.42, d = 2.17· Sep 2026
- RetractedSelf-selection as a requirement for the reputation experiments, read as corroborating emergent motivation· Sep 2026
- Superseded“Exactly eleven” scripts import the GCL package, with 09 and 10 named as importing nothing· Sep 2026
What currently survives
Six buckets, strongest evidence first. Nothing here is new — every row links to the section that shows the work — but the page is long enough that the status of a claim is easy to lose, and the two things most worth knowing early are which results come from prompted LLMs and which come from scripts that never import the GCL package.
Real LLM experiment
Prompted Claude and GPT agents on arithmetic and counting tasks. The only evidence here that is not simulation.
- External assessment is better calibrated than self-confidence for all four agents — Brier 0.281 vs 0.484 for the worst-calibrated agent; holds 4/4, though one margin is thin (0.334 vs 0.361)
- Self-selection routed 39 of 60 tasks to the most overconfident agent — and still matched external assignment on success rate, 0.700 vs 0.717
Measured in simulation
A real stochastic outcome inside a self-designed environment. Says nothing about the world outside that environment.
- GCL's failure-first contracts reduce hold-ups 40.4% against incomplete contracts — the one Hart–Moore prediction that is measured rather than set; direction is built in, magnitude is not
- GCL sends 23.3% fewer messages than Contract Net — 84.0 ± 47.0 vs 109.5 ± 69.7 across 5 trials — wide, overlapping spreads — while placing last of six on efficiency
- Commitment entropy falls 48.7% over 5,000 timesteps without central coordination — single seed, on code with no test coverage
- Difficulty-weighted reputation cuts the easy-task ratio by 59.8 points — 0.667 → 0.069 over 10 seeds, but it also lowers cooperation 0.382 → 0.337, and the mechanism rewards difficulty directly
- Redemption lifts cooperation past the no-consequences control — 0.319 → 0.619 against a 0.519 control, n = 5 seeds; one of the few results that does run the GCL package
Imposed by construction
The simulation was built so this had to come out. Reported because the magnitude is informative, not because the direction was discovered.
- Incomplete contracts produce under-investment (Hart–Moore P1, P2) — investment is a hardcoded multiplier per condition; the test reads back 0.5 × {1.0, 0.7, 0.4, 0.85}
- Harsher consequences reduce cooperation (r = −0.972) — the severity parameter is defined as the amount subtracted from cooperation probability; the sign cannot come out otherwise
- Self-selection matches a perfect-information oracle under fixed effort — the corrected oracle reduces to the same argmax; the difference is exactly 0.000 in 2,000 of 2,000 draws
- Strategic reputation awareness causes gaming (Experiment 25, p = 7e-41) — the gaming counter only exists inside the strategic branch, and incrementing it does not change which task the agent takes
Null result
Tested with adequate power and did not appear. Kept because it constrains the theory.
- The simulated motivation effect did not transfer to LLM agents — 240 paired tasks, +0.017, McNemar p = 0.56, 95% CI [−0.028, +0.058] — the interval excludes the simulated +0.065
- Template replicator dynamics failed all three sub-checks — fitness/usage correlation is −0.299 — the wrong sign, and significant
- Peer observation alone does not reduce gaming — easy-task ratio moves 0.002, p = 0.48; the experiment records it as effective: false
- In the reputation environment, self-selection did not beat assignment — Experiment 26b: peer assignment won at all four awareness levels, and on effort too — a second simulation that fails to reproduce the motivation effect
- Reputation awareness did not increase cooperation or inequality — both went the other way; three of Experiment 25's four hypotheses failed and the fourth is untestable as written
Retracted or superseded
Published here, then found to be wrong. The original wording and the defect are both kept.
- Template sharing: +17% cooperation, Gini 0.35 → 0.25, bottom quartile 34% faster — traced to hardcoded constants in the chart component; no such numbers exist in the repository
- Four emergent-network figures, including 82.3% protocol-diversity reduction — drawn from hand-picked Gaussians, not from a simulation
- 97% of MARL performance at 25–50× sample efficiency — cherry-picked from a four-baseline grid that also contains 1.00× and 1.20×
- Efficiency halves at ~100 agents; Dunbar-like coordination limit — a grid-resolution artifact — 100 was one of seven sizes chosen in advance
- Self-selection beats a truly optimal oracle via information advantage — the oracle optimised the wrong objective; found by an adversarial re-audit and corrected by a new experiment
- Gaming rates of 12% / 45% / 18% / 15%, and t = 8.42, d = 2.17 — chart constants again; two of the four condition names name no experiment, and the two statistics appear in no artifact
Not tested
Stated on this page as motivation or design rationale. No experiment here bears on it.
- Template sharing transfers capability without leaking private agent information — no private state, no adversary, no leakage metric anywhere in the model
- GCL supports scalable oversight or constitutional-AI-style guardrails — a motivating application, argued from the framework's design rather than from any result
- Templates capture structural similarity in the sense analogical-reasoning research means — the inspiration is real; it is not evidence, and Experiment 24's templates have no structure
One cross-cutting caveat that does not fit a bucket: thirteen experiment scripts — 01, 02, 03, 06, 07, 08, 09, 10, 12, 15, 16, 17 and 18 — import src/gcl/ and measure the framework. That set covers 07 (emergent properties), 08 (message economy), 16 (punishment paradox) and 17 (redemption). Everything from 19 onward defines its own agent classes instead — 21, 23, 24, 36, 37, 39 and 40 — so results from those are results about a small standalone heuristic that shares a name with the framework. A further five (11, 11-v2, 20, 41, 41b) import gcl.llm.api_clients only as API plumbing. See Limitations.
The Punishment Paradox
r = -0.818 (n=5 seeds)
Verified Sep 28, 2026 against commit b2d2e27 of the real simulation code, not a copy of this page's numbers. CI run · source · reproduce full scale in Colab
The finding: in this simulation, increasing the consequences for commitment violations decreases cooperation, monotonically and steeply — 0.501 at no consequences down to 0.231 at full severity.
That runs counter to a simple deterrence intuition, the one that says raising the cost of defection should raise cooperation. It is not a refutation of game theory, and an earlier version of this page said it was. Repeated-game models have described punishment that backfires — through counter-punishment, feuding, and antisocial punishment — for decades, so a negative slope here is unsurprising on the literature's terms as well.
The direction is built into the model. AblationAgent.decide_cooperation() computes a cooperation probability and then, when the partner is under penalty, executes coop_prob −= config.violation_penalty. The swept parameter labelled “consequence severity” is the amount subtracted from the probability of cooperating. There is no channel in this model through which severity could deter anything, so r = -0.972 is substantially a readback of a subtraction. What the simulation does add is the size of the collapse: the direct subtraction is linear, and the observed fall is not, because penalised agents defect, which penalises their partners, which spreads. That amplification is the part worth taking seriously, and it is the part the section below describes.
Don't trust the curve — generate it. Drag a penalty slider and watch cooperation collapse in a live port of the simulation.
Run it yourself →Why It Happens: Retaliation Cascades
High consequences trigger retaliation cascades: penalties cause counter-defection, which spreads through the population. The correlation is strong: r = -0.972, p < 0.001.
What the sweep reports
- • No consequences vs full: t = 52.10, p < 0.001
- • Effect size: Cohen's d = 13.45
- • Monotonic decrease across all 5 levels
- • n = 30 seeds per condition
d = 13.45 is large enough to be a flag, not a flex: it means the simulation's punishment mechanic is close to deterministic, not that real-world punishment effects are this dramatic. Read it as an internal comparison inside a self-designed environment, not an externally validated effect size.
The curve is also scale-sensitive. The 5-seed configuration CI runs on every push produces a much sharper collapse (0.492 → 0.059) and a weaker correlation (r = −0.818) than the 30-seed numbers quoted here. CI reproduces the mechanism; it does not reproduce these figures.
Reproduce this result: see experiments/derive_real_headline_stats.py in github.com/jstiltner/gcl, which runs the real 16_consequence_severity_sweep.py simulation directly. (These numbers were corrected 2026-09-01: the previous values here traced to a synthetic-data generator, not this simulation — see the repo's CHANGELOG.)
Redemption Removes the Paradox in Simulation
The candidate fix: add a redemption pathway that lets agents recover from failures, keeping the penalty in place while shortening how long it is carried. In this environment that reverses the paradox. Whether it does so for the stated reason — that recoverability reduces the fear which suppresses commitment-making — is an interpretation of the result, not something the sweep measures.
In Experiment 17, a 20% redemption boost lifts cooperation from 0.319 under standard consequences to 0.619 — and past the 0.519 reached by removing consequences altogether. That last comparison is the one that matters: redemption does not merely dodge the punishment paradox by softening penalties, it beats switching penalties off. A 30% boost does slightly better still (0.649). 50 agents, 100 rounds, n = 5 seeds — a small sweep, and the numbers should be read accordingly.
Superseded· Sep 2026history
This chart previously showed 39.3% → 60.0%, annotated +52.7%. Those values were not simulation output: they matched 22_statistical_significance.py's test_redemption_mechanism(), which draws 0.35 + N(0, 0.08) and 0.60 + N(0, 0.08) and then runs a real paired t-test on the invented samples. The source repo's CHANGELOG identified that generator on 2026-09-01 and listed this claim as one it had not yet re-verified. Re-verified now against the real Experiment 17 run, the effect is larger (+94.0%, not +52.7%) on fewer seeds (5, not 30). The claim survived; the number did not. The chart's error bars, previously labelled “CI”, were and are standard deviations.
How Redemption Works
- 1.Failed agents can attempt recovery actions
- 2.Successful recovery reduces permanent reputation damage
- 3.Effort costs prevent gaming
- 4.Order effects controlled via eligibility snapshots
Hart-Moore Patterns Reproduced (Experiment 21)
42.9% hold-up reduction (n=5 seeds)
Verified Sep 28, 2026 against commit b2d2e27 of the real simulation code, not a copy of this page's numbers. CI run · source · reproduce full scale in Colab
GCL's failure-first contracts are motivated by Hart-Moore incomplete contract theory (Nobel Prize 2016). Experiment 21 reproduces the four patterns that theory describes inside a purpose-built simulation. That is a consistency check on the model, not external evidence for Hart-Moore and not evidence that GCL inherits its results — the theory is the independently established thing here, and the simulation is the thing being checked against it.
The four predictions are not equally informative, and reporting them as “4 of 4 confirmed” obscures that. Two are parameter readbacks, one is a structural separation, and one is a genuinely measured outcome. They are labelled individually below.
Prediction 1: Complete Contracts Enable Investment
by constructionAgents with complete contracts invested 0.500 against 0.200 under high-incompleteness contracts, on this model's 0–1 investment scale.
This is a parameter readback, not a result. investment_decision() sets investment to 0.5 × multiplier with the multiplier hardcoded at 1.0, 0.7, 0.4 and 0.85 for the complete, low-incompleteness, high-incompleteness and GCL conditions. The observed means across 30 seeds are 0.5001, 0.3501, 0.2000 and 0.4252 — the constants, recovered to four decimal places. No t or d is quoted because the statistic here (d = 514) measures only how little noise was added.
Prediction 2: GCL Approaches Complete Contract Benefits
by constructionGCL agents invested 0.425 against 0.200 — 75% of the complete-contract benefit over that baseline.
Same readback as Prediction 1, and the number that matters is the one nobody measured: 0.85 was chosen. Whether GCL-style contracts would in fact sustain 85% of complete-contract investment is the question this prediction appears to answer and does not.
Prediction 3: Incomplete Contracts Enable Hold-ups
structural separationHigh-incompleteness environments averaged 89.7 hold-up incidents per 30-seed run; complete-contract environments had zero. Hold-ups are structurally impossible when every contingency is pre-specified, so t = 66.26 and d = 17.11 separate a real stochastic count from a real structural zero. Both sides are honest; the gap between them was never in doubt.
Prediction 4: GCL Reduces Hold-up Vulnerability
measuredGCL reduced hold-up incidents by 40.4% (95% CI: [37.2%, 43.5%], bootstrapped; t = 19.74, d = 5.10) versus the high-incompleteness condition — 89.7 down to 53.5.
The only one of the four with a stochastic outcome on both sides. Its direction is still designed in: negative contingencies carry a 6× higher hold-up probability than positive ones, and the GCL condition is defined as the one that pre-specifies exactly the negative contingencies. Its magnitude is not, and is the more interesting number — GCL achieves the reduction while carrying more than twice the relationship-specific investment (0.425 vs 0.200), which the same model makes it correspondingly more vulnerable to being held up over.
Corrected 2026-09-01: every number above now comes directly from running IncompleteContractEnvironment, the real agent-based simulation in this experiment — n=30 seeds, 20 agents, 200 timesteps per condition. The previous version of this section showed different numbers (36.8% hold-up reduction, t/d values up to 7.33) that traced to a synthetic-data generator in a separate script, not to this simulation; see the repo's CHANGELOG for the full story. Two further things a reader should know: the effect sizes above (up to d = 17.11) come from a self-designed simulation with a deliberately clean experimental separation and are not estimates of effect size in deployed systems; and 21_incomplete_contract_theory.py does not import src/gcl/ — its “GCL” condition is a set of constants inside the script, so this section describes the framework's design logic rather than the framework's code. Reproduce: experiments/derive_real_headline_stats.py in github.com/jstiltner/gcl.
Coordination Scaling (Experiment 23) — Mostly Retracted
Retracted· Sep 2026history
Experiment 23 does not run GCL. It imports nothing from src/gcl/. Line 35 adds that directory to the path and then never uses it. What the script actually runs is a self-contained NumPy model: Dirichlet capability vectors, argmax selection, and a Bernoulli draw on capability match. Whatever it shows, it is not a property of the commitment-learning system this page is about, and this section presented it as one from the day it was published until September 2026.
Two of the three headline numbers were worse than mislabelled — they were circular. The log-shaped decay was an input to the model, not a finding of it.
The claim here used to be that coordination efficiency degrades logarithmically with population size (R² = 0.88, p = 0.0017), implying a Dunbar-like limit around 100 agents. The fit is arithmetically correct and the plotted data below is the real output of the script. Neither fact makes the conclusion a result.
Scaling Results — what each one really is
- • “~100 agents: efficiency drops to 50% of maximum” — retracted, grid resolution. Peak efficiency is 0.265 at ten agents, so half-max is 0.1325. The sweep tests 5, 10, 20, 50, 100, 150 and 200; n=50 sits at 0.184 and n=100 at 0.114. The crossing is somewhere between them, and 100 is simply the first point sampled below it. The answer could only ever have been one of seven numbers chosen in advance.
- • “Messages grow at 0.08 per agent (sublinear)” — retracted, circular. Line 331 sets messages = int(np.log(n + 1) * 5 + Poisson(3)). Sublinear growth was written into the generator and then reported as though it had been observed.
- • Task concentration: Gini rises 0.354 → 0.980 — stands, with a caveat. This one is not assumed anywhere; it is a real consequence of repeatedly argmax-ing over fixed capabilities. But it is a property of the toy model, not of GCL.
Network Topology — retracted in full
All three figures previously in this card have no counterpart in results/23_dunbar_scaling/results.json. They came from a hand-written summary block, not from the run.
- •
Clustering: low global density (0.12)— actual value 0.0 - •
Small-world: not detected (coefficient 0.89)— actual value 0.0 - •
Structure: hub-and-spoke topology emerges— no topology at all
Mean degree is 0.0 at every population size but one. The trust matrix initialises as np.eye(n) * 0.5 and an edge requires trust above 0.5, which nothing in the run reaches — so the graph being measured had no edges in 69 of 70 runs. The third bullet described the shape of a network that did not exist.
Also not a finding: the results file reports a phase transition at a population of 5. That is the smallest size tested. It is returned because coordination efficiency never crosses the 0.5 threshold the script looks for at any population — the highest value anywhere in the sweep is 0.265 — so the search falls through to its first candidate. It was never published on this page, and it is named here because the same defect produced two of the numbers that were.
Clarification: Task concentration (Gini) measures how tasks distribute across agents—larger populations concentrate tasks on fewer high-reputation agents. This differs from role specialization (HHI < 0.02 in Experiments 33-34), which measures whether agents focus on specific task types. The two can diverge: an agent may handle many tasks without specializing in any particular type.
Emergent Network Properties
Retracted· Sep 2026history
This section previously reported four figures — 82.3% protocol-diversity reduction (R² = 0.91), trust clustering 0.699 at global density 0.12, a task-concentration Gini of 0.745, and a 26.5% efficiency improvement (r = 0.73) — “across 1200+ runs, all p < 0.001”. Those numbers did not come from a simulation. They match 22_statistical_significance.py's test_population_dynamics() to ten decimal places, and that function draws from four hand-picked Gaussians (clustering = 0.7 + N(0, 0.1), gini = 0.75 + N(0, 0.08), and so on) before running genuine t-tests on the invented samples. The p-values were real; the data were not.
This is the same generator the repo's CHANGELOG caught on 2026-09-01 behind the punishment-paradox and Hart-Moore figures. That entry fixed those two claims and listed this one as an open item it had not yet re-verified. It is now re-verified. The real Experiment 07 numbers are below, including the prediction that failed.
Experiment 07 runs a 100-agent population for 5,000 timesteps and tests four pre-registered predictions about what structure should emerge without explicit coordination rules. Three passed, one failed. Single seed (seed=42) — this is one run, not a distribution.
Schematic. This diagram is procedurally generated to illustrate hub-and-spoke topology; it is not a plot of the Experiment 07 trust graph.
Protocol Convergence passed
Agents converge on shared commitment templates without central coordination.
Commitment entropy falls 2.696 → 1.384 over 5,000 timesteps (48.7%), fitting a power law with α = 0.165, R² = 0.782.
Trust Network Clustering passed
Trust edges close into triangles rather than staying tree-like.
Clustering coefficient 0.7475, mean path length 1.33, 4,248 edges. The network is dense(density 0.429) — the earlier “sparse … density 0.12” framing was backwards as well as unsourced.
Specialization passed
Agents concentrate on narrower slices of the task space over time.
Specialization index rises 0.223 → 0.755 between the early and late windows of the run.
Template Replicator Dynamics failed
Predicted: successful templates should out-replicate unsuccessful ones.
All three sub-checks failed. Fitness/usage correlation is −0.299 — the wrong sign, and significant (p = 3.1e‑20). 909 of 1,000 templates stayed in use; dominance ratio 0.40 against a predicted concentration.
Retracted· Sep 2026history
This box used to describe an unresolved disagreement. There was no disagreement. It said Experiment 07 reported trust clustering of 0.7475 while Experiment 23 — the Dunbar scaling sweep below — reported exactly 0.0 at every population size, and that the two could not both be right. Chased down in September 2026: Experiment 23 was not measuring a sparse network. Its trust matrix initialises as np.eye(n) * 0.5, leaving every off-diagonal entry at zero against an edge threshold of > 0.5 that nothing in the run ever crossed. The graph had no edges at all in 69 of 70 runs.
The correction does not vindicate the figure above. It removes the only thing that was ever offered as a check on it. Experiment 07's 0.7475 is a single seed, on code with no test coverage, and now nothing else in the repository measures the same quantity — so the honest status of GCL's trust topology is no supported claim, not a confirmed one. Details in Limitations.
Source: results/07_population/results.json (predictions block), produced by experiments/07_population_dynamics.py. Note that src/gcl/population/, the code behind this experiment, has no test coverage — see Limitations.
The GCL Framework
Where the two mechanisms came from. This work started from a long-standing worry about linguistic representation. Rather than assume agents share semantics, or try to solve representation directly, it asks a narrower question: can coordination be grounded in externally observable consequences instead? That produced grounded commitments — coordinate around observable success and failure conditions rather than inferred intentions — and templates, from the role analogy plays in human learning.
Both began as cross-domain hypotheses and were then made testable. That origin is hypothesis generation. It is not evidence for either mechanism, and nothing below should be read as support drawn from it. What the mechanisms are worth is decided by the experiments, which is why several of them are marked retracted.
Why This Formalization?
Traditional multi-agent coordination assumes agents share semantic understanding. GCL replaces this assumption with verifiable behavioral contracts:
- • Trigger (τ): When does this commitment activate? Removes ambiguity about scope.
- • Action (a): What behavior is promised? Observable, not interpretive.
- • Verification (φ): How do we know it succeeded? Third-party verifiable.
- • Failures (F): What can go wrong, and what happens then? Enumerated, not implicit.
- • Stake (σ): What does the agent risk? Skin in the game.
Formal Definition
A grounded commitment is a 5-tuple:
- — trigger predicate
- — action function
- — verification predicate
- — failure modes with stakes and remediations
- — stake (reputation at risk)
1[COMMITMENT]
2ISSUER: Agent_A
3TRIGGER: Task requires capability X
4BEHAVIOR: Complete subtask within 3 rounds
5SUCCESS: Subtask verified complete
6FAILURES:
7 - IF timeout THEN stake_loss=0.5, REMEDIATION: delegate
8 - IF capability_mismatch THEN stake_loss=0.2, REMEDIATION: escalate
9 - IF resource_exhaustion THEN stake_loss=0.3, REMEDIATION: request_resources
10CONFIDENCE: 85%
11STAKE: 1.0
12[/COMMITMENT]Design choice: failure-first
Unlike traditional contracts that specify success conditions, GCL commitments enumerate failure modes, and success is the complement of all of them. The argument for this is auditability: an auditor knows what to check, and an agent has fewer undefined edge cases to claim success in. That is the intent behind the formalization. It is not a property anything on this page demonstrates — no experiment here puts an adversary in front of a commitment and measures whether it can be satisfied without being met.
Commitment-Grounded Learning
Agents learn what commitments to make via reinforcement learning. The policy maps states to commitment portfolios, optimizing for expected value minus stake risk. The sketch below is written to be read, not to be run — there is no GroundedCommitmentLearner class in the repository. The real implementation is split across gcl/learning/: CommitmentEnv, CommitmentPolicy, RewardComputer and PPOTrainer.
1class GroundedCommitmentLearner:
2 """Agent that learns what commitments to make via reinforcement learning.
3
4 Key insight: Agents don't need shared understanding, just shared consequences.
5 """
6
7 def __init__(self, capabilities, stake_budget):
8 self.capabilities = capabilities
9 self.stake_budget = stake_budget
10 self.reputation = ReputationTracker()
11 self.template_library = TemplateHierarchy()
12
13 def propose_commitment(self, task, context):
14 """Policy maps states to commitment portfolios."""
15 capability_match = self.assess_capability(task)
16 observability = self.assess_verifiability(task)
17
18 if capability_match < 0.5 or observability < 0.3:
19 return None # Refuse rather than risk failure
20
21 failure_modes = self.enumerate_failures(task)
22
23 return Commitment(
24 issuer=self.id,
25 trigger=task.trigger,
26 behavior=task.required_behavior,
27 success=task.success_condition,
28 failures=failure_modes,
29 confidence=capability_match * observability,
30 stake=self.calculate_stake(expected_value, risk)
31 )Template Sharing (Experiment 24)
Simulation— standalone abstract simulation; does not run the GCL packageRetracted· Sep 2026history
This section previously reported four figures — directed sharing +17% cooperation (t = 4.5, p = 0.001), bottom quartile improving 34% faster than top quartile, inequality falling Gini 0.35 → 0.25, and random sharing at +8%. None of those numbers appear anywhere in the GCL repository. They came from the figure below, whose bar heights and 200-round curves are hardcoded constants in TemplateSharingChart.tsx — the “+17%” is 0.61/0.52, two of those constants divided by each other, and the chart draws that string on the canvas as an annotation. The prose then quoted the annotation as a measurement.
The real artifact, results/experiment_24_template_sharing.json, reports directed sharing at +5.2 percentage points over no sharing (0.555 → 0.607, t = 13.11) and a Gini of 0.066 → 0.000, not 0.35 → 0.25. It also contradicts the ordering the old copy asserted: random sharing finishes at 0.6013 and directed at 0.6008. On the artifact's own headline metric, directed sharing is not the best policy. The corrected reading is below, and it is much weaker than what it replaces.
One further caveat on that artifact, found 2026-09-20: the file is truncated and is not valid JSON. It stops at 1,886 bytes in the middle of the key predictions.pred1_cooperation_improved.direction_correct, so the run was interrupted before it finished writing and the experiment's own pass/fail verdicts were never recorded. The comparison figures quoted above come from the statistical_tests block, which was written before the truncation and is intact — but nothing here should be read as a completed experiment.
The hypothesis is that agents should transfer reusable structural patterns rather than relearn every coordination problem from scratch. Experiment 24 is the toy test of that hypothesis. It is worth being precise about what it does and does not show, because the answer is mostly not much.
Cognitive inspiration — hypothesis generation, not evidence. Templates came from an analogy to analogical learning: humans who learn that "promising to deliver X by deadline Y" works in one context reuse the pattern elsewhere instead of rediscovering it. The adjacent literature is structure mapping (Gentner, 1983) and case-based reasoning(Kolodner, 1992).
That is where the mechanism came from. It is not why anyone should believe it works. Nothing on this page tests whether GCL templates capture structural similarity in the sense Gentner means, and Experiment 24 in particular does not: its "templates" carry no structure at all.
Schematic, retained to show what was retracted. Every series in this figure is generated in the browser from start + (end − start)(1 − e^−3t) + noise with hand-chosen endpoints. It is not a plot of Experiment 24. The corrected numbers are in the table below.
Experimental Conditions
- • No Sharing: Templates never transfer (baseline)
- • Random Sharing: Random pairs exchange at 10% rate
- • Directed Sharing: High → low capability transfer
- • Mutual Sharing: Bidirectional exchange
What the artifact reports
| Policy | Coop. | Gini |
|---|---|---|
| No sharing | 0.5609 | 0.054 |
| Random | 0.6013 | 0.000 |
| Directed | 0.6008 | 0.000 |
| Mutual | 0.5986 | 0.000 |
10 runs × 50 agents × 200 rounds; mean of the last 20 rounds.
Three reasons to discount this experiment
- The ceiling is 0.60, and every sharing condition sits on it. An agent succeeds with probability
capability × (1 − 0.8·difficulty), capability is clamped at 1.0, and difficulty is drawn uniformly from [0.3, 0.7] — so the maximum attainable cooperation rate is exactly 0.600. Random, directed and mutual land at 0.6013, 0.6008 and 0.5986. They are indistinguishable because they are all at the analytic ceiling, which is also why Gini reaches exactly 0.000: everyone hits the clamp. The “inequality reduction” is the clamp, not an effect. - “Cooperation” is not cooperation. The metric is
round_successes / n_agents— the fraction of agents that individually passed a solo task. Agents in this simulation never interact, never coordinate, and make no commitments. Nothing in the file importssrc/gcl/; it defines its own three-fieldAgent, and a “template” is a float between 0.05 and 0.15 added to a capability score. It has no structure to map. - The statistics and the artifact are both unsound. The reported t = 13.11 treats 50 consecutive rounds of a single averaged trajectory as 50 independent samples. The properly clustered test across the 10 runs was computed but never reached disk: the committed JSON is truncated mid-write at 1,886 bytes and is not valid JSON. Seeds are set from
hash(condition), which Python salts per process, so the run is not reproducible either.
Design implication, not a finding. Sharing a template transfers what worked without transmitting the recipient-visible internals of the agent that found it. That is a reason the mechanism is worth building and testing. It is not something this experiment measured: there is no private state in the model, no adversary, and no leakage metric of any kind. The earlier claim that template sharing “provides a mechanism for capability transfer without revealing private agent information” is a hypothesis about a system nobody has built yet.
Current status of the template mechanism: not yet tested. The cold-start behaviour in the MAS comparison below is the closest thing to evidence that reusable structure helps, and it runs a different codebase.
Why This Might Matter for AI Safety
Qualified· Sep 2026history
Nothing in this section is a result. It used to read as one. The earlier wording said GCL “provides a foundation for verifiable multi-agent coordination” and produces coordination that is “verifiable, auditable, and aligned”. No experiment in the repository tests an oversight property. There is no adversarial agent, no overseer, no attempt to hide a violation from a commitment log, and no measurement of what an auditor can or cannot recover. The strongest claim the evidence supports is narrower and worth stating plainly: behaviour is verifiable against a contract inside the designed framework, because the framework defines success and failure as observable conditions. Whether that survives contact with an agent trying to defeat it is unknown.
The reason to pursue this line of work is an oversight question: how do you verify what agents are doing when you cannot inspect their internal states? Below are three properties the design is intended to have, and what would have to be true for each to hold.
Auditability
Argument: every action traces to a commitment with enumerated failure modes, so an auditor knows what to check, and an agent cannot claim success by exploiting an edge case nobody defined.
Depends on the failure enumeration being adequate. Hart-Moore is a theory about contracts being necessarily incomplete; Experiment 21 shows the payoff when the right contingencies happen to be specified, and says nothing about how often they will be.
Accountability
Argument: the stake parameter (σ) creates skin in the game. Agents that make unreliable commitments lose reputation and future coordination opportunities.
Depends on reputation not being gameable. The gaming-resistance experiments below are the only work bearing on this, and they are simulations against a modelled attacker, not a real one.
Policy constraint
Argument: if agents may only make commitments matching approved templates, the template set becomes a place to express policy at the coordination layer.
Untested, and the weakest of the three. No experiment restricts a template set or checks that the restriction binds. The earlier framing of this as “constitutional AI-style guardrails” implied a demonstrated capability and has been removed.
The idea, stated as an idea
GCL is an attempt to sidestep the interpretation problem rather than solve it. The bet is that agents do not need shared understanding if they have shared consequences — that coordination can be grounded in externally observable success and failure conditions instead of in inferred intent.
That is a research direction with a partial simulation record behind it. It is not a principled foundation for aligned coordination, and this page should not be read as claiming it is. The one place the bet has been tested against real language models — Experiments 41 and 41b — is also the place where half of it failed to reproduce.
Self-Selection vs. External Assignment
Retracted· Aug 2026history
An earlier version of this page reported that self-selection beats optimal external matching by 81%, driven ~75% by information asymmetry. That claim is retracted. The original oracle (Experiment 39) scored candidates by closeness-of-fit, penalizing over-qualified agents under a success model that is monotonically increasing in capability—it was not optimal. Against a corrected, truly success-maximizing oracle (Experiment 40), the information advantage is +0.000. The corrected findings are below.
How this correction happened, and why it is the best result on the page
- Experiment 39 produced a strong, publishable positive result: self-selection beat “optimal” external matching by a wide margin, and a decomposition attributed ~75% of it to agents knowing things a coordinator doesn't. That is what this page reported.
- A stronger model was then asked to falsify the preserved record rather than extend it. It found that Experiment 39's oracle scored candidates by closeness-of-fit — penalising over-qualified agents — under a success model that increases monotonically in capability. The oracle was optimising the wrong objective. It was never optimal, so beating it meant nothing.
- The criticism was not simply accepted. Experiment 40 was written to check it: same environment, corrected argmax oracle, 50 seeds. The claimed information advantage went to exactly +0.000. The retraction below is the result of that run, not of the critique.
- The corrected model made two new predictions — a residual motivation effect of about +0.065, and the claim that external observation should beat self-assessment whenever self-assessment is noisy. Both are testable outside the simulation.
- Experiments 41 and 41b tested them on real LLM agents. The calibration prediction transferred: external assessment was better calibrated than self-confidence for all four agents. The motivation mechanism did not: a powered paired design found nothing, with a 95% CI tight enough to exclude the predicted effect (240 paired tasks, +0.017, p = 0.56, CI [−0.028, +0.058] vs a predicted +0.065).
A positive result became a retraction, a null, and one prediction that survived contact with a different substrate. The resulting account is less favorable and better supported, and it is the sequence this page should be judged on.
Corrected finding: With effort held fixed, self-selection and a truly optimal (argmax) oracle are statistically indistinguishable. The surviving mechanism is emergent motivation: agents that choose their own tasks develop higher effort, beating the optimal oracle by +0.065 cooperation (95% CI [+0.050, +0.080], d = 1.68). Choice creates commitment—not privileged information.
Scope of that sentence: it describes Experiment 40's simulation and nothing else. The same effect was looked for in real LLM agents and not found (41b, 240 paired tasks, p = 0.56, with a 95% CI that excludes +0.065 — detailed below), and it reversed in a second simulation (26b, where assignment beat choice at all four awareness levels). Nothing here demonstrates emergent motivation in an LLM.
Corrected Mechanism Decomposition
With effort fixed and perfect information on both sides, self-selection (0.531) and the corrected argmax oracle (0.531) are identical: +0.000 [−0.013, +0.013], d = 0.00. The previously reported +0.24 gap is fully reproduced as the gap between the corrected and flawed oracles.
This null is analytic, not measured. Self-selection scores volunteers by perceived − difficulty × 0.4, and that second term is the same for every agent, so the rule is argmax over perceived capability. With zero noise on both sides, the oracle is argmax over the same quantity — the two are the same function, and pick the same agent in 2,000 of 2,000 draws. Fifty seeds could not have returned anything else. The ±0.013 interval is not a residual effect; it is the two runs' random streams drifting apart when nobody volunteers and self-selection falls through to a random draw. Reported here because the confidence interval and the d = 0.00 imply a test that could have come out differently. It could not. That makes the null stronger than a measured one, not weaker — it is not a question of statistical power — but it is less work than the statistics suggest.
Self-selection with emergent effort (0.599) beats the optimal oracle (0.534) by +0.065 (d = 1.68). Assigned agents do not develop the same commitment dynamics (oracle emergent ≈ oracle fixed: +0.003).
The Observability Phase Boundary
Sweeping agent self-knowledge noise (σself) against coordinator observation noise (σoracle) shows neither mechanism intrinsically dominates:
- →Self-selection wins iff σoracle > σself (advantage up to +0.127)
- →Central assignment wins iff σself > σoracle (symmetrically, down to −0.127)
- →Design principle: delegate task selection to whichever party has the less noisy view of agent capability, plus a motivation bonus to self-selection worth roughly one noise level
The same caveat applies here. Both mechanisms are argmax over a noisy estimate of the same underlying scalar, so “whoever sees more clearly wins” follows from how the model is built rather than from running it — and the exact symmetry of the two bounds (+0.127 against −0.127) is imposed by that construction, not discovered in it. The sweep quantifies the effect; it does not establish the boundary. What the simulation genuinely produces is the third row: the +0.065 motivation effect is a feedback dynamic, not an identity, and it is the only part of this section that could have come out at zero.
Taking it to real LLMs (Experiments 41 / 41b) — one prediction held, one did not
The phase boundary makes a testable prediction: self-selection helps real LLM systems only if models assess their own success probability better than an external observer can. Testing with Claude and GPT agents on verifiable tasks:
- →No privileged self-knowledge: external assessment was better calibrated than self-confidence for 4/4 agents (e.g., Brier 0.281 vs 0.484)—current LLMs sit on the central-assignment side of the boundary
- →Self-selection rewards overconfidence: pure self-selection routed 39/60 tasks to an overconfident agent (0.60 accuracy) over a perfectly reliable one (1.00 accuracy)
- →Motivation effect not detected in LLMs: a powered paired re-test found no significant "volunteered vs assigned" framing effect. Pooled over both agents (240 paired tasks): +0.017, McNemar p = 0.56, exact 95% CI [−0.028, +0.058] — an interval that excludes the simulated +0.065, so this is a null with the power to rule the predicted effect out, not merely a failure to find it. Per agent: claude-reasoning +0.000 (p = 1.0), gpt-reasoning +0.033 (p = 0.52). The CIs are derived from the artifact's discordant-pair counts; the artifact itself reports only the point estimates and p-values.
Experimental Design
- • Self-select vs corrected argmax oracle (Exp 40)
- • Fixed effort (0.8) vs emergent effort conditions
- • 5×5 observability noise sweep, 50 seeds/cell
- • LLM calibration + framing tests (Exp 41/41b)
What the runs report
- • Motivation effect: +0.065 [+0.050, +0.080], d = 1.68
- • Information effect: +0.000 [−0.013, +0.013] (null)
- • Phase boundary significant when noise gap ≥ 0.1
- • LLM framing null: McNemar p = 0.56, 240 paired tasks
Design Implications
Whether to use self-selection is an empirical question about relative observability, not a universal principle. In this simulation choice creates commitment. It is the only environment in this repository where it does: the effect is absent in prompted LLM agents (41b) and reversed in the reputation environment (26b, where assignment won at all four awareness levels). Two out of three, and the two negatives are the ones closer to a deployed system.
The one implication that survives all three: for current LLM agents, self-reports are less reliable than external assessment, so routing should be grounded in verified track records rather than agent self-assessment. Grounding reputation in observable outcomes rather than stated confidence is what the GCL formalization is for — though no experiment here tests a routing system built that way.
Full correction history and canonical claims table: github.com/jstiltner/gcl.
Gaming-Resistant Reputation Mechanisms
What was measured. In Experiment 27, difficulty-weighted reputation cut the easy-task ratio from 0.667 to 0.069 — a drop of 59.8 percentage points, p = 7e-34, 10 seeds per condition. Here easy_task_ratio is the share of tasks agents actually chose from the easy end of the difficulty range, so it is a behavioural measure. It still measures selection, not intent, and nothing checks whether the easy tasks were the wrong ones to pick.
Experiment 25's version of this number is not a measurement. It reports an easy-task ratio of 0.966 under strategic awareness against 0.000 under every other condition, p = 7e-41. The counter behind it (easy_tasks_selected) sits inside an if awareness_level == "strategic" branch, so it is structurally unable to increment anywhere else — the three zeros are unreachable code, not observed restraint. Worse, incrementing it does not change which task the agent takes; only the commitment stake is adjusted. The counter tallies how often a low-reputation agent was offered an easy task. That p-value is reported here only because it was on the page before, and it should be read as describing the task distribution rather than any agent behaviour.
Retracted· Sep 2026history
This section previously reported four conditions — Blind Reputation 12%, Naive Visible 45%, Difficulty-Weighted 18%, Social + Difficulty 15% — and headlined “reduces gaming behavior by 59.8% (t = 8.42, p < 0.001, d = 2.17)”. None of those four rates appear in any result artifact. They were constants in GamingReductionChart.tsx. Two of the condition names do not exist in any experiment: there is no “Social + Difficulty” run and no “Naive Visible” run.
t = 8.42 and d = 2.17 do not appear in any artifact either, and the scripts do not compute them — Experiment 27 reports a p-value and a mean difference and nothing else. Those two statistics have no traceable source and have been removed rather than recomputed.
The 59.8 was real but mislabelled. It is the absolute change in easy-task ratio (0.667 → 0.069) recorded in the artifact's gaming_change field, not a 59.8% relative reduction. Relative, the reduction is 89.6%. The chart above now plots all six of Experiment 27's real conditions.
The winning mechanism has a cost
Difficulty weighting is the most effective anti-gaming mechanism of the six tested, and it also lowered cooperation — from 0.382 to 0.337, p = 4e-06. Every mechanism that significantly reduced gaming also moved cooperation down or left it unchanged; none improved both. Agents push into harder tasks and fail more of them.
Social pressure alone did nothing. Peer observation without difficulty weighting changed the easy-task ratio by 0.002 (p = 0.48) and the cooperation rate by 0.003 (p = 0.65). The experiment records it as effective: false. It is a clean null and it is the mechanism closest to what a deployed peer-review scheme would actually be.
Experiment 25 tested four hypotheses and none survives
- N/AStrategic awareness increases gaming — recorded as the one pass, but the gaming counter only exists in the strategic branch, so the comparison is unavailable as a test. See above.
- FAILAwareness increases cooperation — it decreased it, significantly and in the wrong direction: 0.271 blind → 0.184 self-aware, p = 2e-09.
- FAILAwareness increases inequality — Gini fell slightly with awareness (0.771 blind → 0.701 strategic) rather than rising.
- FAILSelf-awareness without social visibility is optimal — the best condition on the experiment's own composite score was blind, i.e. no reputation visibility at all.
In this environment, assignment beat self-selection
This box used to claim these experiments ran voluntary task selection throughout, and to read them as corroborating the emergent-motivation result from Experiment 40. Neither is right. Experiment 26b runs the contrast directly — four awareness levels crossed with peer assignment versus free choice, 10 seeds — and assignment won in all four:
- blind assigned 0.2767 · choice 0.2749
- self assigned 0.2719 · choice 0.2664
- social assigned 0.3022 · choice 0.2793
- strategic assigned 0.3049 · choice 0.2650 (p = 0.004)
Mean effort went the same way — 0.87 under strategic-assigned against 0.84 under strategic-choice — so the motivation channel Experiment 40 isolates does not show up here at all. Experiment 26 separately ranks peer_assignment as its best mechanism of six. These are different environments with different task models, so this is not a refutation of Experiment 40; it is a second simulation, built by the same author, that fails to reproduce its direction. Taken with Experiment 41b's null in real LLM agents, the motivation effect should be read as a result about one environment rather than a property of self-selection.
What the sweep reports
Easy-task ratio: 0.667 → 0.069, p = 7e-34
Cooperation: 0.382 → 0.337, p = 4e-06
10 seeds per condition. The scripts report p-values and mean differences; they do not compute t-statistics or effect sizes, so none are quoted.
Mechanism
Difficulty weighting divides reputation gain by task difficulty, so easy tasks return less credit. That the easy-task ratio falls is close to definitional; what is not definitional is how far it falls and what it costs in cooperation.
Design implication, not a finding
If a deployed system scores agents on a metric agents can see, the metric should price difficulty. That is a suggestion carried over from this simulation, not something any experiment here tested outside it.
GCL vs. Multi-Agent RL
Qualified· Sep 2026history
The cooperation-rate comparisons below (Experiments 36 and 37 — the “third of five” ranking, both tables, and the non-stationary results) do not run src/gcl/. Neither script imports it. Each defines its own standalone class called GCLAgent / GCLSelfSelection — the same ~15-line volunteer-if-capable-enough heuristic, copy-pasted between the two files, structurally identical to Experiment 40's self_selection function. 36_marl_comparison even imports Agent, Task and create_population from the project's agent harness and never calls any of them. None of GCL's calculus, verification, protocol or reputation machinery runs in either experiment. The chart below is unaffected: Experiment 08 does import gcl.multiagent and gcl.baselines, so its message-overhead and efficiency numbers are the real package. Nothing here is arithmetically wrong — the 0.534 and the tables below are correctly transcribed from their scripts' own output — but “GCL” in these two experiments names a threshold rule, not this repository.
Key finding: Experiment 36 ranks its self-selection heuristic third of five on final cooperation — behind IQL (0.553) and QMIX (0.542), ahead of MAPPO (0.522) and random (0.475). The results file records gcl_is_best: false. All four learned methods sit inside each other's 95% confidence intervals; only the gap to random is a real effect (d = 1.10). 30 agents, 2,000 episodes, 20 seeds.
Qualified· Sep 2026history
This section previously led with “GCL achieves 25–50× better sample efficiency than MARL while maintaining 97% of MARL's coordination quality”. Both halves were true of a chosen subset. Experiment 36 reports sample efficiency against four baselines:
“25–50×” was the range across the two favourable baselines with the other two dropped. MAPPO is a MARL algorithm and converges in exactly the same 2.05 episodes GCL does, so “MARL requires 52–102 episodes” was false as stated. The 97% figure compared GCL only to IQL, the single best baseline, which reads as a narrow shortfall; it omits that QMIX also beats GCL and that GCL ranks third overall.
The efficiency ratios are also unstable. episodes_to_50 for QMIX is 102.4 with a standard deviation of 435.4 and a CI lower bound of 0.8; for IQL it is 51.75 ± 214.6, CI lower bound 0.8. In most seeds the baselines cross the threshold immediately too — the means are carried by a few outlier runs. And the threshold itself is 50% cooperation, which the random baseline reaches 0.475 of on its own.
What survives is a difference in kind rather than a measured speedup: the heuristic reaches its asymptote in ~2 episodes because it does not learn — it reads a capability value handed to it rather than discovering one through trial, the way MARL does. That is a real property of the volunteer-threshold design, not evidence about this repository's implementation of it, and it is why GCL is intended to be usable cold-start. It is not a like-for-like efficiency comparison, and Experiment 36 shows the design does not buy better final coordination.
Retracted· Sep 2026history
The chart above previously plotted invented coordinates: CNP at 142 messages, FIPA-ACL at 156, Auction at 128, and MARL at 168, against a y-axis labelled “Task Success Rate” on which GCL scored 0.87. Only GCL's message count (84) came from Experiment 08. Every other value was fabricated, and every error ran the same direction: competitors' message counts inflated, their scores pushed below GCL's. The real values are CNP 109.5, FIPA-ACL 190.2, Auction 175.6, and MARL-IQL 0 — it is a non-communicating method and sends no messages at all. On mean_efficiency, the metric the experiment actually compares (success_rate is 1.0 for all six and does not discriminate), GCL is last of the six at 0.645 and is strictly dominated by MARL-IQL, which sends fewer messages and scores higher. GCL is not on the Pareto frontier. Its genuine result is the message count: 23.3% fewer than CNP, per the experiment's own key_findings block — not the “~40%” the chart annotated.
Episodes to 50% Cooperation
Mean ± SD over 20 seeds. The SDs are larger than the means.
Final Cooperation
Mean over 20 seeds. GCL ranks third; the top four CIs all overlap.
Non-Stationary Environments
Experiment 37 (experiment_37_regime_changes.json), difficulty regime shifts, 30 agents, 10 seeds. Change frequency is episodes between shifts, so 50 is the most volatile column and 1000 the least.
| Change every | 50 | 100 | 200 | 500 | 1000 |
|---|---|---|---|---|---|
| GCL − IQL | +0.61 | +0.90 | +0.21 | +0.27 | +0.17 |
| GCL − QMIX | +0.74 | +1.38 | +1.33 | +1.54 | +2.57 |
Percentage points of mean cooperation. Per-arm SDs are 1.0–2.3 points at n = 10, so every cell here is inside noise.
Retracted· Sep 2026history
This block previously showed four buckets — Low / Medium / High / Very High volatility at +0.2% / +0.5% / +0.7% / +0.9% — under the heading “GCL's advantage grows in non-stationary environments”. Experiment 37 has five change frequencies, not four buckets, and neither of its two series matches those values. The ordering was also inverted: change frequency counts episodes between shifts, so a lower number means a more volatile environment. Read correctly, the two baselines disagree about the direction of the effect. Against IQL the advantage is largest in the volatile conditions (+0.90 points at frequency 100, +0.17 at 1000), which is the predicted pattern. Against QMIX it runs the other way: +0.74 at the most volatile setting rising monotonically to +2.57 at the most stable one — the opposite of the claim. crossover_iql and crossover_qmix are both null: no crossover point was found. With ten seeds and per-arm SDs of 1.0–2.3 points, none of these deltas is distinguishable from zero, and the “advantage grows” claim is not supported either way.
When to Use Each Approach
GCL is preferred when:
- • Cold-start coordination (no training data available at all)
- • Privacy-preserving systems (agents keep capabilities private)
- • Agents already hold reliable self-knowledge to read
MARL may be better when:
- • Training time is available
- • Final coordination quality matters — IQL and QMIX both beat GCL here
- • Agent capabilities are fully observable
“Non-stationary environments” was previously listed on the GCL side and has been removed: Experiment 37's two baselines disagree on the direction of that effect and all of its deltas are within noise. Sample efficiency was also listed and is now qualified above — MAPPO matches GCL episode for episode.
Limitations
On method. AI was used heavily throughout this work — hypothesis exploration, implementation, experimental iteration. It is not what makes any of the above true. Where a model proposed an explanation for a result, that explanation was treated as the next hypothesis to test, not as the finding.
This note previously read “every number on this page traces to an executable experiment with a controlled comparison, a statistical test, and a stated seed count.” That was not true when it was written. Six figures on this page did not trace to an experiment at all, and the sentence asserting they did was itself doing rhetorical work. It has been replaced with the accounting below rather than quietly softened, because a claim of universal provenance is exactly the kind of claim that should not be made on assurance.
Six figures previously on this page were not simulation output, and one section described a resolved bug as an unresolved disagreement — see the research record for the full correction history. The forward-looking limitations below are unaffected by those corrections.
The Population Code Has No Test Coverage
The suite is 382 passing tests at 63% line coverage overall, but src/gcl/population/ — every module behind Experiment 07, roughly 1,350 statements — is at 0%, as are core/scope.py and llm/api_clients.py. The emergent-properties results rest on code that nothing exercises except the experiment script itself.
GCL Does Not Win on Coordination Quality
Experiment 36 ranks GCL third of five methods on final cooperation, behind IQL and QMIX, and records gcl_is_best: false. The four learned methods' confidence intervals all overlap, so the honest reading is that GCL is competitive rather than better. Its distinguishing properties are cold-start operation and message economy, not coordination quality.
Simulation Environment
Every result here except Experiments 41 and 41b was produced in a simulation designed alongside the hypothesis it tests. Deployment would introduce network latency, partial observability and adversarial agents, none of which appear in any experiment on this page.
The Later Experiments Do Not Run the GCL Package
Superseded· Sep 2026history
This box previously read “exactly eleven” scripts and named 09 and 10 as examples that “fall inside that numeric range but import nothing from src/gcl/”. That was wrong. Both import the commitment calculus — from src.gcl.core.calculus import Commitment, FailureMode, Outcome, issue, verify, settle, confidence — and both call it; 10_cicd_coordination.py runs the full issue/verify/settle lifecycle. The correct count is thirteen.
The defect was a grep that missed a spelling: ^from gcl does not match from src.gcl.core.calculus. Worth keeping because of its direction — the error understated how much of this work runs the real package, and an error that makes a page look more self-critical attracts less scrutiny than a flattering one. An inaccurate disclosure is still an inaccuracy.
Thirteen of the experiment scripts import the GCL framework: 01, 02, 03, 06, 07, 08, 09, 10, 12, 15, 16, 17 and 18. Those supply the emergent-properties, message-economy, punishment-paradox and redemption results. Within that range only 04, 05, 13 and 14 define their own agent classes instead. The break comes after 18: everything from 19 onward is standalone, including 21 (Hart–Moore), 23 (Dunbar scaling), 24 (template sharing), 36 (MARL comparison), 37, 39 and 40 (self-selection).
A further five scripts — 11, 11-v2, 20, 41 and 41b — import gcl.llm.api_clients and nothing else. That module is plumbing for reaching the model APIs, so those runs are real experiments on real LLMs but are not exercises of the coordination framework either.
Results from the standalone scripts describe a small heuristic that shares a name with the framework, and should not be read as measurements of the framework's implementation. The thirteen above do measure the package — but 07, 12, 15, 16, 17 and 18 all run through src/gcl/population/, which has no test coverage at all.
Task Complexity
Experimental tasks are simplified relative to production multi-agent systems. Commitment verification in complex, multi-step tasks may require additional mechanisms not yet validated.
No Known Scaling Bound
This entry used to assert a coordination limit around ~100 agents and recommend hierarchical structures above it. That bound was retracted — see Dunbar scaling above. The figure was a grid artifact: the sweep tested 5, 10, 20, 50, 100, 150 and 200, and 100 was simply the first sampled size below half of peak efficiency. The largest population anyone has run here is 200 agents, on a script that does not import src/gcl/. The honest position is that GCL's scaling behaviour is unmeasured, not that it is flat up to 100 and hierarchical beyond it.
Motivation Effect Is Simulation-Only
The emergent motivation effect (+0.065, d = 1.68) was measured in simulation but was not detected in prompted LLM agents: a powered paired re-test (Experiment 41b, 240 paired tasks over two agents, McNemar) found no significant "volunteered vs assigned" framing effect (+0.017, p = 0.56), with a 95% CI of [−0.028, +0.058] that excludes the +0.065 the simulation predicted. LLM self-assessments were also less calibrated than external assessment, so self-selection results should not be assumed to transfer to LLM systems.
Ongoing Work
The findings above come from a numbered experiment series running to 41b, and from ten corrections made in place — including the 39 → 40 → 41 → 41b chain, where a positive result was falsified, re-measured, and then failed to reproduce in real LLM agents. Current focus areas:
Papers in Preparation
- →Choice Creates Commitment: Emergent Motivation in Self-Selected Coordination
In preparation. The title overstates what is currently in hand: the effect is one simulation's result that did not reproduce in prompted LLM agents or in a second simulation. Any submission has to lead with that.
- →When Should Agents Choose Their Own Tasks? An Observability Phase Boundary
In preparation for peer review
Future Directions
- •Mapping the observability phase boundary across agent architectures
- •Measuring where coordination actually degrades with population size. This previously read “hierarchical GCL for populations > 100 agents”, which assumed a bound that has since been retracted as a grid artifact.
- •Any deployment at all. Every result here except Experiments 41 and 41b is simulation.
- •Testing whether observable success conditions can carry any part of a constitutional-AI-style constraint — currently an argument, not a result.
Open Questions
The corrected findings raise a deeper question: why does choosing a task change how hard an agent works on it?
- ?Is there a motivation channel at all?
Experiment 40 finds self-selecting agents developing higher effort than assigned agents (+0.065, d = 1.68) while agents given identical tasks do not. It is the only place that effect appears. It is absent in prompted LLM agents (41b, McNemar p = 0.56 over 240 paired tasks, with a CI excluding +0.065) and reversed in the reputation environment (26b, assignment ahead at all four awareness levels, and on effort too). The open question is not yet what mechanism produces the effect; it is whether the effect is a property of choice or a property of Experiment 40's task model.
- ?Commitment templates as geometric operations
Can template learning be characterized as geometric transformations on a pre-linguistic capability manifold? If so, template sharing may be transferring operations rather than knowledge. This is speculation with no experiment behind it, and the one experiment that shared templates represented a template as a float added to a capability score, so it has no structure to bear on the question either way.
- ?Emergent coordination primitives
What computational primitives underlie the emergent specialization we observe? Experiment 07's trust clustering (0.7475) is a single seed on code with no test coverage, and nothing else in the repository independently measures the same quantity, so this is a question to re-run before it is a question to explain.
Explore the Implementation
Interested in discussing this work? Email me