GMAA Governed Multi-Agent Architecture

Field note

The Second Empty Chair: Why Your AI's Context Window Is an Ungoverned System

Session degradation in AI agents is not a capacity problem. It is an ungoverned commit process at the context boundary: everything commits, nothing is authorized, and no one answers for what is kept or dropped.

The first empty chair: The Coherence Boundary. · After publishing: what seven models said.

I don't build transformer models. I run multi-agent Claude Code systems that do real work, pipelines with dozens of subagents writing to shared state, and I kept hitting the same wall every operator hits: the session degrades. The first hour is sharp. By the third, the model contradicts decisions it made earlier, relitigates settled questions, and loses the thread it was hired to hold.

So I went digging into the mechanics. Why does context fill up? And why can't the model just use something like an index server that retrieves what it needs, when it needs it, instead of dragging the whole transcript around?

The answers were straightforward and unsatisfying. A transformer reprocesses the entire conversation on every response. Attention is all-to-all, so cost grows quadratically with length. Nothing persists between turns; the session is an illusion the chat application maintains by resending everything. And the index server exists. It's called retrieval, and every memory vendor sells a version of it. But retrieval scores chunks one at a time against a query. It cannot see that two chunks jointly contradict each other, or jointly imply something neither says alone.

That last sentence stopped me. Per-item evaluation missing what only exists at the set level. I had written that argument already. It's my own thesis, applied to a boundary I hadn't pointed it at.

The window is a coherence boundary

GMAA, the framework I authored at gmaa.ai, makes one core claim: the unit of authorization is the set of pending changes, not only each change alone. A coherence boundary is a defined set of shared state whose invariants must hold together. Everyone proposes into it, one writer commits, and the complete pending set is ratified by an accountable human before anything lands. At the code boundary that seat is a person. At the boundary this article is about, it will not be, and the reason is the argument. I built it for multi-agent systems writing to shared repositories and databases, because individually correct changes were breaking systems jointly, and nobody was looking at the set.

Now look at what a context window actually is. It is the complete working state of a reasoning system: every commitment made, every constraint accepted, every fact established across the session. The model's next output is conditioned on all of it, jointly. If two entries in the window contradict each other, the reasoning built on them is incoherent, in exactly the way a migration and a query change that each pass review can jointly break a production system.

Here is what that looks like in practice. Early in a session, the operator confirms the staging database is disposable and safe to drop. An hour later, a tool result notes that staging now serves the client demo. Each entry is individually fine. Each was committed into the window without review against the other. A model holding both will confidently execute a drop that was authorized against a world that no longer exists. No single entry is wrong. The set is.

The context window is a coherence boundary. It may be the most consequential coherence boundary in any AI-assisted operation. And in every shipping system, it has no writer, no set, and no seat.

Everything commits, nothing is authorized

Watch what enters the window in an ordinary agentic session: every token of every draft, every tool result including the failed ones, every exploration that led nowhere, every subagent's output as it arrives. Each item lands the moment it is produced. No entry is ever reviewed against the entries already there. The window is an append-only log where commit authority belongs to whoever writes fastest. In other words, it belongs to no one.

This is per-change commit without even per-change review. It is the configuration GMAA exists to name as a failure mode, and it runs at the heart of every AI system in production. My session degradation wasn't a capacity problem. It was an authorization problem wearing a capacity problem's clothes.

The consequences are documented. Chroma's technical report gave the degradation its name, context rot, after testing 18 frontier models and finding that reliability drops as input grows, even on simple tasks, and even when every relevant fact remains retrievable.1 The pollution itself costs capability. A model reasoning over its own accumulated exhaust reasons worse, whether or not the signal is still in there somewhere.

The industry built two-thirds of the answer

I checked whether anyone had solved this before claiming the gap. The serious players have stopped ignoring it. Anthropic ships compaction as a core mechanism in Claude Code and publishes context engineering as a discipline.2 Claude Code's own permission prompts are the pattern in miniature: per-action approval, asked and answered one action at a time, with no moment where the set is put to anyone. Letta treats the window as virtual memory, paging state in and out like an operating system, and deserves specific credit: its memory blocks are human-inspectable and agent-editable, which is closer to governed state than anything else shipping. The memory-layer vendors, including Mem0, Zep, LangGraph, and AWS AgentCore, run extraction and consolidation pipelines. Some show good audit-trail instincts, such as marking old entries invalid rather than deleting them.

Map these against GMAA's three moves and a pattern emerges. Move one, individual review, is everywhere: relevance scoring, recency weighting, LLM judges grading what to keep. Move two, assembling and checking the set, exists in fragments. The paging architectures and DAG-based state proposals at least treat context as structured state rather than a transcript. Move three exists nowhere in formalized form: no shipping system places an independent, accountable seat over the complete set with authority to ratify or halt before commit. Inspectability is not authorization. A memory block a human can read is not a regime a human has ratified and answers for.

In every system I could find, the entity actually deciding what commits into the boundary is an automated evaluator sharing substrate with the agent it serves. Compaction is the model summarizing its own history. Extraction is a pipeline judging items one at a time. The research community's own characterization work found that summarization output barely responds to prompting,3 which means today's retention process is unaccountable and unsteerable at the same time. The operator cannot reliably dictate what survives, and no one answers for what doesn't.

When compaction silently drops the constraint that turns out to matter three sessions later, there is no ratified regime to audit and no accountable party to ask. There is a prompt template.

What governance looks like at this boundary

Applying GMAA here does not require modifying it. The window is a coherence boundary, and the framework already specifies how boundaries are governed.

The draft, meaning the deliberation, the tool noise, and the dead ends, lives in a working space and never enters the boundary. What commits is the ratified set: the decision reached, the state resolved, the facts established as durable. The context stops being a recording of how the work happened and becomes a ledger of what the work concluded. Version control, not a keylogger.

The seat works the way it must wherever human judgment cannot act at machine speed: the accountable human ratifies the regime, not each event. To be explicit, because readers default otherwise: the seat is an office, not a person. At this boundary the office is filled by machinery. The machinery acts under a regime a human wrote and can revoke. The human is not in the loop per set. The human is above the loop and owns the rules the loop runs. In operation, the control loop is simple. The human authors a retention constitution once: what classes of content earn permanent residency in the boundary, what must halt for live review, what never commits. Machinery enforces it per set, at machine speed. The human sees two things: a periodic audit briefing of what committed and what was dropped under the constitution, and live exceptions, the sets the constitution flags as exceeding what it can safely auto-ratify. Those wait for the human or never land. Standing revocation authority completes the loop; the person who ratified the regime can stop it at any moment. This is how aviation, trading, and nuclear operations resolved the same constraint decades ago. Authorization moved earlier. It never moved to the machine.

What this solves, and what it honestly doesn't

Governed commit attacks the pollution problem, which is the dominant practical reason my sessions, and yours, degrade. Signal density rises, fill rate falls, and retrieval improves because what gets reloaded later is coherence-checked state rather than raw transcript that may jointly contradict itself.

It does not touch the capacity problem. A governed window still pays quadratic attention cost on whatever it holds, and state that must be reasoned over jointly must still sit in the window together. The capacity side has its own serious research program: longer-context architectures, hierarchical memory, agentic decomposition. That work is real and advancing, and none of it substitutes for governance, because the context-rot finding cuts the other way: bigger windows do not fix pollution, they give it more room. Governance never claimed jurisdiction over the ceiling, and I won't inflate the claim: the context problem was always two problems wearing one name. Ungoverned commit was the fixable half.

One residual must be owned rather than hidden: ratification is compression, and compression discards. Some detail judged non-durable will matter later. No constitution eliminates that class of loss. A constitution bounds it, makes it auditable, and makes one person answerable for the bounds. That is the same result GMAA delivers everywhere else, and it is the difference between a system that forgets by policy and a system that forgets by accident.

The second chair

GMAA's founding observation was that every multi-agent platform approves changes and none ratifies sets. An empty chair at the commit boundary. The context window is the second chair, and it is empty in every system on the market: the one boundary whose coherence conditions everything the model does next, governed today by heuristics no one wrote down and no one answers for.

I went looking for the mechanics of why my sessions degrade. I found my own thesis, sitting at a boundary nobody had pointed it at. The field is converging on the vocabulary of budgets, curation, boundaries, and invalidation. It has not arrived at the conclusion, because the conclusion is not a better algorithm. It is a seat.

What seven models said about this. I gave the piece to seven models and asked each what it thought. None challenged the core claim once the seat was read correctly, all seven demanded a benchmark, and six of seven read the seat as a person. Read the reactions.

Questions

Why do long AI sessions degrade?

Because everything commits into the context window the moment it is produced: drafts, failed tool calls, dead ends, subagent noise. No entry is reviewed against what is already there, so individually fine entries accumulate into jointly incoherent state. Research calls the result context rot: reliability drops as input grows even when every relevant fact remains retrievable.

Is the context window a governance problem or a capacity problem?

Both, but the fixable half is governance. The window is a coherence boundary, a set of shared state whose invariants must hold together, and today it has no writer, no set-level check, and no accountable seat over what commits. The capacity half, quadratic attention cost, is architectural and falls only to different mathematics. Bigger windows do not fix pollution; they give it more room.

What would a governed context window look like?

The draft, meaning deliberation, tool noise, and dead ends, lives in a working space and never enters the boundary. What commits is the ratified set: decisions reached, state resolved, facts established as durable. An accountable human ratifies the retention regime once; machinery enforces it per set at machine speed; flagged exceptions wait for the human or never land.

Does this fix the quadratic cost of attention?

No, and it does not claim to. A governed window still pays full attention cost on whatever it holds; state that must be reasoned over jointly must still sit together. Governance controls what enters the window; attention cost is a function of how much is present. The two are orthogonal, which is also why the approach needs no change to the model: it removes pollution, not the ceiling.

Is this a replacement for compaction or retrieval memory?

It is a layer above them, not a substitute. Compaction and retrieval evaluate items one at a time; the missing move is set-level authorization on top of that per-item review, never instead of it. Under a governed boundary, compaction and retrieval still run, but what they operate on is ratified state rather than raw transcript, and an accountable human owns the regime that decides what survives.

Who ratifies the set, a human or an agent?

The seat is an office. What fills it depends on the boundary. At the context boundary, a human writes the retention regime once and owns it. Machinery then ratifies each pending set against that regime at machine speed. Only the exceptions the regime cannot settle go to the human. The per-set ratifier here is machinery working under a human-owned regime. It is not a person approving every set. Two things never move to the machine: writing the regime and the standing power to revoke it. Inspectability is not authorization. A memory a human can read is not a regime a human has ratified and answers for.

1Hong, Troynikov, Huber. "Context Rot: How Increasing Input Tokens Impacts LLM Performance." Chroma, July 2025. research.trychroma.com/context-rot ↩

2Anthropic. "Effective Context Engineering for AI Agents." anthropic.com/engineering. Compaction documentation: platform.claude.com/docs ↩

3Cim, Topcu, Das, Kandemir. "Parallel Context Compaction for Long-Horizon LLM Agent Serving." Pennsylvania State University. arxiv.org/abs/2605.23296: summary volume is essentially input- and prompt-invariant; the operator has no fine-grained control since prompt instructions are largely ignored. ↩

The GMAA specification (v1.6.1) is at the architecture; the founding argument is the thesis, The Coherence Boundary. Licensed CC BY 4.0. Canonical: gmaa.ai. Cite by version.