<!-- generated by scripts/generate-agent-docs.ts -->

> **Aegis**
> Infrastructure layer for autonomous agents: durable state, verifiable commitments, policy enforcement. Complements agent frameworks like LangChain/LangGraph. Single-node, 303 tests, no performance benchmarks yet.
>
> Source: https://jasonstiltner.com/projects/aegis/

---

# Aegis

Infrastructure Layer for Autonomous Agents

Updated Sep 8, 2026

A systems-architecture problem, not a model problem: agent commitments as first-class, event-sourced objects; a policy gateway that enforces constraints at tool-invocation time, not after the fact; durability and recovery strategies for failures that happen mid-workflow, not just at the edges.

This is not a "better LangChain." Agent frameworks handle orchestration—what should the agent do next? Aegis handles infrastructure—how do we ensure commitments are kept across failures? You can run a LangGraph workflow on Aegis to gain durability and verification that LangGraph doesn't provide natively.

*The commitment model is built on [Grounded Commitment Learning](https://jasonstiltner.com/projects/grounded-commitment-learning/)'s verifiable-behavior contracts — the research gave the infrastructure its object model, not the other way around.*

Event-sourced state · GCL commitments · Policy gateway · 303 tests

Status: single-node, 303 tests passing, no performance benchmarks yet and no production deployment. The state, commitment, policy and LLM layers are real, tested code. The recovery, multi-agent, built-in-tool and MCP layers are scaffolded but untested (0% coverage), and several of their handlers return hard-coded success. See [Limitations](#limitations).

[View on GitHub](https://github.com/jstiltner/aegis)

## The Gap

#### Workflow Engines (Temporal, Prefect)

State survives restarts. Tasks retry. Execution history is clear. But no agent-specific abstractions: no commitments, no policy evaluation before tool invocation, no multi-agent coordination primitives.

#### Agent Frameworks (LangChain, AutoGPT)

Tool calling, memory, planning. But state is ephemeral—restart the process and you lose everything. No checkpoint/replay, no formal verification of what the agent promised, no structured recovery.

#### What's Missing

Neither treats agent commitments as first-class objects. When an agent says "I will complete this task by 5pm," that's a string in a conversation. No mechanism to verify fulfillment, detect violation, or recover gracefully.

## Architecture

[Diagram: Aegis layered architecture diagram showing 5 layers: Interface, Runtime Core, Tool Gateway, Commitment & Verification, and Audit & Tracing]

Requests flow top-to-bottom; events flow bottom-to-top. The state machine coordinates: receives events, applies transitions, triggers checkpoints, notifies listeners.

## Core Components

### 1\. Event-Sourced State

State derives from an append-only event log. Each user message, LLM response, tool invocation, and commitment update is an event. Current state computes by replaying events from the last checkpoint.

aegis/core/state.py

```python
class AgentState(BaseModel):
    model_config = ConfigDict(frozen=True)

    def with_message(self, message: Message) -> AgentState:
        return self.model_copy(
            update={
                "conversation_history": (*self.conversation_history, message),
                "version": self.version + 1,
            }
        )
```

**Trade-off:** Storage grows with event count; replay adds latency on restore. Configurable checkpoint intervals mitigate this—checkpoint every N transitions or M seconds, whichever comes first.

### 2\. Commitments as First-Class Objects

Commitments use the GCL 5-tuple:`(debtor, creditor, action, condition, deadline)`. The condition field contains an evaluable expression, not a description.

aegis/commitments/models.py

```python
class RuntimeCommitment(BaseModel):
    model_config = ConfigDict(frozen=True)

    debtor: str      # Who made the commitment
    creditor: str    # Who receives the commitment
    action: str      # What was committed
    condition: str   # Evaluable expression: "task_complete AND error_count == 0"
    deadline: datetime | None
    status: CommitmentStatus  # CREATED → ACTIVE → FULFILLED/VIOLATED/CANCELLED
```

This enables: verification (check if condition holds), violation detection (deadline passed, condition failed), recovery (select and execute strategy), audit (track commitment lifecycle).

### 3\. Policy Enforcement at the Gateway

Policy enforces at invocation time, not planning time. The gateway sees actual arguments—an agent might plan to "read a file" but the actual path could be `/etc/passwd`.

aegis/tools/policy.py

```python
class PolicyRule(BaseModel):
    name: str
    action: PolicyAction  # ALLOW, DENY, REQUIRE_APPROVAL
    tool_pattern: str     # Glob: "file_*", "web_search"
    argument_conditions: dict[str, Any]  # {"path": {"not_contains": "/etc"}}
```

**Trade-off:** Can't prevent the agent from wasting tokens planning a disallowed action. The cost of a rejected tool call is low compared to the security benefit.

## GCL Integration

GCL provides the theoretical foundation; Aegis provides the runtime. The GCL 5-tuple maps directly to `RuntimeCommitment`:

| GCL Concept | Aegis Implementation |
| --- | --- |
| Debtor | `commitment.debtor` (agent ID) |
| Creditor | `commitment.creditor` (user/system ID) |
| Action | `commitment.action` (string) |
| Condition | `commitment.condition` (evaluable expression) |
| Deadline | `commitment.deadline` (datetime) |

When GCL isn't installed, Aegis falls back to its own expression evaluator. Supports basic comparisons, logical operators, and membership tests. Unsafe expressions (function calls, imports, attribute access) are rejected.

## Validation

303 unit tests, all passing (`pytest`, 2026-09-08, Python 3.13). They live in a single flat `tests/unit/`directory; the breakdown below is the actual per-file count, not an estimate.

#### test\_llm.py

Client, streaming, providers: 68 tests

#### test\_tools.py

Gateway, policy, auth: 48 tests

#### test\_audit.py

Event log, audit trail: 46 tests

#### test\_api.py

FastAPI routes and schemas: 37 tests

#### test\_gcl.py

GCL integration, verification: 35 tests

#### test\_state.py

Event-sourced state: 24 tests

#### test\_checkpoint.py

Checkpoint integrity: 23 tests

#### test\_state\_machine.py

Transitions, validation: 22 tests

#### What 303 tests does not cover

Line coverage over `src/aegis` is **41%** (4,449 of 8,096 statements unexecuted). Four subsystems have no tests at all and sit at 0%:

-   `recovery/` — detector, orchestrator, strategies
-   `multiagent/` — messaging, patterns, protocol, registry
-   `tools/builtin/` — filesystem, web, code, system
-   `tools/mcp_adapter.py` — the MCP transport layer

`core/replay.py` is at 26% — there is no dedicated replay test. Coverage is highest where the object model is: `state.py` 97%, `gcl/models.py` 95%, `events.py` 92%.

## Limitations

#### Single-node only

State stores locally. Distributed coordination (multiple agents across nodes, shared state) requires a distributed event log (Kafka, Redis Streams) and consensus for checkpoint coordination. Not implemented.

#### No content-aware policy

Constitutional AI principles check metadata, not content. Evaluating whether a response "contains harmful content" requires an external classifier.

#### Recovery strategies do not yet execute a recovery

The strategy selection logic — classify a violation, pick a plan, order it by priority — is real. The execution step is not. `RetryStrategy.execute()` sleeps for its backoff delay and returns a hard-coded `{"success": True}` above the comment *“in a real implementation, this would re-execute the action”*; `RenegotiateStrategy.execute()` does the same without contacting the creditor. Nothing in `recovery/` is covered by a test.

#### The HTTP API and CLI are surface, not implementation

The FastAPI schemas and routing are tested (37 tests), but `POST /sessions/{id}/messages` returns a literal `"This is a placeholder response."` rather than invoking an agent, and the CLI's status and replay commands render placeholder data. The tests verify the contract, not that anything is behind it.

#### No CI, and no property-based tests

The 303 figure comes from running the suite locally; the repository has no CI workflow, so unlike the [CNL benchmark](https://jasonstiltner.com/projects/collaborative-nested-learning/) there is no reproducible artifact behind it. An earlier version of this page also claimed property-based tests via Hypothesis and multi-agent message-ordering tests; neither exists — Hypothesis is a declared dev dependency that is never imported.

#### No performance benchmarks

Checkpoint latency, message throughput, and policy evaluation overhead have not been systematically measured under realistic workloads.

## Explore

[← GCL Framework](https://jasonstiltner.com/projects/grounded-commitment-learning/) · [Full Corpus](https://jasonstiltner.com/corpus/)
