Field note · short version r4 · August 24, 2026
Governed Context: The Short Version
The complete design is at gmaa.ai/governed-context.html. That page is the reference, carrying the full mechanism, every named failure mode, and the review history that changed the mechanism rather than the rhetoric. It runs about 7,600 words because reviewers need that depth. This is the argument itself, in about 1,600 words, measured.
The diagnosis
The context window is a coherence boundary. Everything an agent produces lands there the moment it's produced, no entry is reviewed against what's already present, and individually acceptable entries accumulate into jointly incoherent state. Research named the effect context rot, where reliability drops as input grows even when every relevant fact is still retrievable. A major and under-addressed source of that decline is not capacity but ungoverned commit, which is the half that governance can reach.
The claim, and the test it has to pass
How long a session stays sharp depends on what its memory is allowed to hold. Right now it holds everything, including the dead ends and the answers you already replaced. Under this design it holds only what the current work still needs. What's left agrees with itself, because anything new has to say what it replaces. The old version retires right then, instead of sitting there looking just as current.
The acceptance criterion is stated before any measurement exists, and it's two-sided. The regime has to show no capability regression against today's baseline, and it has to keep a session useful much longer before it hits a wall. Both have to hold or the design fails its own test. Nothing has been benchmarked. This is a design hypothesis with serious groundwork behind it, not a solved problem.
Who governs
The seat at this boundary is machinery under a human-owned regime. A seat is a position of authority defined by what it may decide, not by who fills it, and this one is filled by a rule set executing. A person, the operator, authors the retention constitution once and versions it over time, naming what earns residency, what retires, and what is ambiguous enough to stop and ask. The machinery, the gate, enforces it per set at machine speed, and it's the only thing that commits to the window. When a case falls outside the rules, the system halts to the operator instead of guessing, because halting on ambiguity is what keeps the seat free of agent judgment even though no person sits in it. The operator holds audit and revocation and doesn't act at commit.
This governance is additive. Set-level authorization sits on top of whatever per-item evaluation a system already runs. It never replaces it.
Where the seat actually goes
The first objection a working engineer raises is that none of this is theirs to govern. The window is a runtime artifact of the serving stack. The app resends a transcript, the harness compacts it when it fills, and every rule about what gets kept was set by the model's developer in code nobody downstream can see or amend. On that reading the accountable seat has nowhere to sit, and the design is a request to vendors rather than something anyone can adopt.
That reading is right about the API and wrong about where the window comes from. Something assembles the context before every call, and today that something is an agent framework, an orchestrator, or a harness somebody on your side operates. Whatever it sends is the window. So the gate goes there, one layer above the model, and the person accountable is whoever runs the system rather than whoever trained the weights. No change to the transformer is required, and none is asked for. If the lab ever governs admission inside the serving stack, that's better still, but nothing here waits on it.
The three edges
Governed commit, at entry. Scratch work stays scratch. Only ratified conclusions enter, the decisions made, the facts established, the questions resolved. Every admission declares what it supersedes, so replacement retires the old entry instead of leaving two live-looking answers. And admissions are authorized as sets rather than items, because two entries that are each acceptable alone can be incoherent together, and no item-level check can see a property that exists only between items.
The compaction czar, in the interior. Ratified state still accumulates, so something has to decide what stays resident, and the earlier answer, keep it until superseded, produced retention inflation. The fix is two graphs. Provenance is permanent and records what was derived from what. Residency is separate and records what must remain resident for future reasoning. Only residency controls the window, and residency is computed rather than declared. The constitution names the roots, current goals, unresolved constraints, active decisions, open obligations, and an object stays only if it's reachable from a live root. Edges carry policies and leases, so a defensive dependency that's soft and unrenewed expires. The default flips from stay-until-superseded to stay-while-justified. None of this puts judgment in the czar, since roots are declared classes, reachability is traversal, and leases are a clock. The czar proposes compaction sets, the constitution ratifies them, and unclear cases halt.
The record, at exit. Nothing is lost. Everything the session produces goes into the record the moment it's produced, including the scratch work and the failed attempts, and nothing there is ever deleted or quietly rewritten. The record and the window are two separate boundaries rather than two shelves in one store. The record's job is to keep everything and mark its status. The window's job is to hold a set that agrees with itself. The gate sits between them, which means retiring something is two moves, a status change in the record and a removal from the window, and nothing ever leaves the record at all.
One rule falls out of that, and it's the sharpest thing in the design. The record is never the judge. If the storage layer reranks or summarizes or otherwise improves itself in the background, it has changed what the model sees without anything authorizing the change, which is the same ungoverned writer this design exists to remove, hiding one floor down. Search over the record can propose. Only the gate disposes, as a set, like any other admission. That also settles an argument the public reviews left open, whether dead ends are pollution or worth keeping. They're neither in the window and both in the record, tagged as tried and failed, costing the window nothing and coming back only through the gate. And since retirement is a status flip rather than a deletion, keeping something is no longer the only way to avoid losing it, because nothing gets lost.
What it does not promise
The regime makes correctness auditable. It doesn't guarantee it. A false declaration, filed correctly, passes, the same way it passes today, with one difference that carries the whole argument. Under governance the false claim is recorded with provenance, what it retired is recallable, and every declared descendant has a traceable edge to it. Under transcript accumulation the same lie is absorbed as prose with nothing marking it at all. A structured mistake can be found. An invisible one can't.
The open problems that will decide it
Set construction above declared targets. Something decides what belongs to a pending set and when it closes, and that's a hidden seat. The design occupies it with a rule the constitution can check. A set is the minimal closed collection of proposals sharing a declared target or dependency edge inside a defined window. Proposals with overlapping but non-identical targets, soft dependencies, and shared assumptions never declared as targets are outside that rule's reach, and above the level of declared targets, set-level authorization is a floor rather than a solution. Named with severity, because it's the make-or-break question for the correctness claim.
Undeclared epistemic descendants. When a set is invalidated, the walk forward through provenance marks every declared descendant at-risk and emits a bounded blast radius. Descendants that never declared their ancestor escape the walk, and moving anything from at-risk back to live is a new epistemic act by a writer, not a mechanical result. Narrowed, not closed.
Halt rates. If the constitution flags too much, the operator is back to per-item review under another name and the regime collapses into noise. This is the central operational risk, so it's made measurable rather than assumed away. Halt rate, halt categories, and resolution time are primary metrics, and if the rate can't be driven low enough for a class of work, the scope claim narrows instead of the detector loosening.
Scope and ontology. The regime claims work whose state has a type system, and open-ended interpretive work is named as future research rather than covered. Even inside scope, where the type system comes from is open. A fixed ontology is too rigid, and an evolving one means ontology changes are themselves governed sets that put dependent state at-risk until revalidated.
The invitation
The full design carries all of this at working depth, plus the failure modes this summary doesn't reach and the six rounds of adversarial review that rebuilt the mechanism. Read it at gmaa.ai/governed-context.html if you intend to build on it, and more carefully if you intend to attack it. Name the failure modes it misses and say plainly where the claims outrun the evidence. That's the whole method.
Questions
Is the seat at the context window a person?
No. That is the misread this page exists to prevent. At the code boundary the governed seat is a human. At the context window the same thesis lands as machinery under a human-owned constitution: the operator authors and versions the rules the gate runs, and no person sits at per-set commit.
Is any of this benchmarked?
No. This is a design draft. The claim made here is narrower: the seat can be occupied today at the assembly layer, where context is already being constructed mechanically. Nothing has been benchmarked, and the open problems above are what a benchmark would have to settle.
Where is the full design?
At gmaa.ai/governed-context.html, about 7,600 words: the complete regime, the failure modes, the review history, and the questions for reviewers. This page is the entry; that page is the reference.