Field note · design draft r9 · August 24, 2026
Governed Context: an end-to-end regime for the AI context window
Review draft r9, 2026-08-24. Israel Heskiel, gmaa.ai. This isn't a solved problem. It's a new direction and serious groundwork toward solving it, and this version says what's open. The design incorporates six rounds of adversarial review from two independent reviewers, and the reviews changed the mechanism rather than the rhetoric. The changes and their reasons are recorded in section 10. The proposal is to change the storage discipline of long-running AI work from append-only transcript accumulation to an explicit, governed lifecycle. The acceptance criterion is no capability regression against today's baseline and a session that runs much longer before it hits a wall. The regime claims work whose state has a type system, and it leaves open-ended interpretive work as future research. This is a design hypothesis, and nothing in it has been benchmarked, so attack the mechanism, name the failure modes it misses, and say plainly where the claims outrun the evidence.
Read this first. This is the complete reference, about 7,600 words at reviewer depth, with the full mechanism, every named failure mode, and the review history. If you want the argument first, the short version is about 1,600 words. Come back here to go deeper, or to attack the mechanism.
Sections 1, 5, and 7 are where the document lives. Section 5 is the failure modes. Section 10 is what I want from you.
Terms
Seat. A seat is a position of authority in the system, and it's defined by what it may decide, not by who or what fills it. A seat isn't a person. Any seat may be filled by a human, by an agent, or by a rule set executing, and which of these fills it is a design decision made per boundary. Where this document says the seat isn't human, it means the position exists but no person occupies it at that point in the flow.
Operator. The operator is the accountable human. The operator authors the retention constitution, holds audit and revocation, and resolves halts, but the operator doesn't act at commit.
Constitution. The constitution is the retention constitution, a rule set the operator authored once and versions over time, and the machinery enforces it per set. At the context boundary the constitution fills the seat at commit. Its computational power is bounded on purpose. Its rules are deterministic predicates over typed declarations and metadata (set type, declared supersession, declared dependencies, provenance, and status), and it may not invoke a model, contain a learned classifier, or perform semantic classification such as deciding that a claim is durable, that two claims conflict, or that a proposed dependency is valid. Any step that would require that kind of classification is a halt, not a rule. That's the line between permitted evaluation and forbidden judgment, and if a deployment crosses it, the constitution has become an agent in the seat under another name.
Czar. The czar is the compaction component. It's a software component and not a role a person holds, so it proposes compaction sets and executes declared supersession without exercising judgment of its own, and where judgment would be needed it halts.
Gate. The gate is the constitution executing at commit. It's the one committer at this boundary. Writers propose into it, and it commits or halts, and nothing else writes to the window.
Set constructor. The set constructor is whatever decides which proposals belong to one pending set and when that set closes. This is a position of real authority, because splitting {A, B} into two sets can produce a different authorization result than considering them together, and the whole premise of set-level governance is that set boundaries matter. This document doesn't yet fully specify the set constructor, as section 5 explains.
Writer. A writer is any process that proposes state into the boundary, whether that's an agent, a tool, or the model itself in a chat setting. Writers propose but never commit. A writer is also a semantic trust boundary, because the gate verifies that declarations are well-formed but can't verify that they're true, so the reliability of the declarer is part of the design's security model rather than outside it.
Set. A set is the unit of authorization. It's the complete group of pending changes considered together, and it includes what each change supersedes and what each change depends on.
1. The problem
Long AI sessions degrade. The model contradicts earlier decisions, relitigates settled questions, and loses the thread. The common diagnosis says the window simply fills, but my diagnosis is different. The window fills through an ungoverned commit process. Every token, tool result, dead end, and subagent output lands in the window the moment it's produced, and no entry is reviewed against what is already there, so individually acceptable entries accumulate into jointly incoherent state. Research calls the effect context rot, where reliability drops as input grows even when every relevant fact is still retrievable.
The public argument for this diagnosis is at gmaa.ai/second-empty-chair.html. This document goes past the argument to the full regime that follows from it.
Two problems share the name "context problem," and this regime addresses only one of them.
Context computation. This is how efficiently a model attends over a given amount of context, from quadratic attention cost to KV cache limits and their remedies. This document doesn't address it.
Context integrity. This is what deserves to become authoritative state, how it reconciles with existing state, and what is presented to the model now. The regime addresses this, and it doesn't change the math of attention. It changes what attention runs over.
The design makes two claims, and they have to be kept apart because they rest on different things.
The primary claim is session lifespan. Governed commit keeps the draft out of the window, and continuous retirement moves superseded state out as later sets land, so window occupancy tracks declared-live state instead of history. Stated precisely, the claim has three parts. The lifecycle transition is mechanical given declarations. Whether those transitions produce a sustained reduction in resident state is the primary hypothesis. Whether the resulting resident state preserves capability is the secondary hypothesis. The mechanism operates under wrong declarations, but whether it still buys lifespan under wrong declarations is what the benchmark measures, and section 5 names the one way it can fail to.
One definitional point belongs here, because it will be raised. Today's baseline has no notion of live at all. The window holds everything until it dies, and occupancy tracks history and only history. So "the window holds declared-live state and retires the rest" isn't a tautology against that baseline. It's the entire change. Declared-live state is what the machinery says must stay resident, while actually-needed state is the minimum required for continued correct performance, and the lifespan hypothesis is that the first stays a useful approximation of the second over long horizons. That gives two errors, under-retention (needed state retired) and over-retention (unneeded state kept), and for the primary claim the second matters more.
The secondary claim is state correctness, that what the window holds is right. That depends on declaration truth, it's where the open problems in section 5 live, and it's the research hypothesis. The regime makes correctness auditable. It doesn't guarantee it.
The acceptance criterion follows from the split. The regime wins if it doesn't make the session worse or less capable than today's baseline, and makes it run much longer before it hits a wall. Those are the two benchmarks named in section 7. The baseline it must not regress from already runs on undeclared topology, already compacts by a model summarizing itself, already loses state silently at capacity, and already dies. Every semantic failure named in section 5 exists in that baseline too, invisibly. So no-regression is the bar, and the design's case for clearing it is that it makes the same class of mistakes declared, retirable, and enumerable instead of hidden. The lifespan question doesn't wait on the correctness result, because the occupancy mechanism runs either way.
One more boundary on scope. The regime assumes that authoritative state can be separated from the information needed to reason over it. That holds for work whose state has a type system, for software projects, program records, and anything with explicit artifacts and supersession semantics. It may not hold for research, investigation, negotiation, diagnosis, or open-ended design, where current state is closer to a distribution over interpretations than to a ledger. Those may be the workloads where degradation matters most, and for them the ontology itself is the open question rather than the mechanism. This document claims the first class and doesn't claim the second.
2. The framework in one paragraph
GMAA makes one claim, which is that the unit of authorization is the set of pending changes, not only each change alone. A coherence boundary is a defined body of shared state whose invariants must hold together. Everyone proposes into it, one writer commits, and the complete pending set is ratified by an accountable human before anything lands. Per-change review is necessary but not sufficient, so set-level authorization stacks on top of it and never replaces it.
The context window is a coherence boundary, because it's the complete working state of a reasoning system and it conditions everything the model does next. In every shipping system a careful survey can find, it has no writer, no set-level check, and no accountable seat over what commits.
One translation is needed at this boundary. The generic GMAA sentence says one writer commits and an accountable human ratifies the set. Here the one committer is the gate, which is the constitution executing, and the accountable human ratified the constitution rather than each set. Writers propose but never commit. The sentence holds, and only the occupants change.
3. The regime: three mechanisms at three edges
The regime governs the window at its entry, its interior, and its exit. Each mechanism is GMAA applied, and none of them changes GMAA.
3.1 Governed commit (entry)
The draft, which is deliberation, tool noise, and dead ends, lives in a working space and never enters the boundary. What commits is the ratified set, the decisions reached, the state resolved, and the facts established as durable. The window is a ledger of what writers successfully declared as concluded and the gate authorized, not a recording of how the work happened.
Each ratified set carries what it decides, what it supersedes, what it depends on, and its provenance. Supersession is declared at ratification, not inferred later.
The ledger isn't only decisions and facts. A conclusion can depend on rejected alternatives, unresolved questions, confidence, negative evidence, and reasoning constraints, and if none of that commits, later reasoning repeats mistakes or reaches a different conclusion from apparently identical premises. So the constitution defines commit classes for deliberative material, and a rejected alternative or an open question is a typed entry that lives, gets superseded, and retires like anything else. This costs some cleanliness at the draft-versus-durable boundary, and that cost is owned rather than hidden. To give the czar a mechanical handle, every deliberative entry carries an expires-after or superseded-by-default declaration, so an open question that nobody renews retires on its lease instead of living forever.
Where the writer is the model itself, as in a chat given a long task, the model proposes supersession as part of the set and the constitution ratifies or halts it. Declared supersession doesn't require a code seat. It requires a task long enough for its own decisions to be revised.
No human acts at commit, because nobody can ratify context updates at machine speed and nobody should have to. The operator authors a retention constitution once, defining what classes of content earn residency in the boundary, what must halt, and what never commits. The constitution, enforced by machinery, ratifies each set. The operator answers for the constitution, sees a periodic audit briefing, and resolves the cases the constitution doesn't cover. Those halt rather than being decided by machinery, and standing revocation authority completes the loop. Accountability is the operator's, while action at the gate is the constitution's.
3.2 The compaction czar (interior)
Ratified state still accumulates, and the question is what must stay resident. An earlier revision answered that with supersession alone, where state stays until something replaces it. That answer produced retention inflation, because it conflates two relations that aren't the same. B was derived from A, and that fact is permanent provenance, but once B is committed, B may no longer need A resident. So the design keeps two graphs. The provenance graph is permanent and records what was derived from what. The residency graph is separate and records what must currently remain resident for future reasoning. The archive preserves history, the provenance graph preserves derivation, and the residency graph alone controls the window.
Residency is then computed, not declared. The constitution names the roots: current goals, unresolved constraints, active decisions, open obligations, active plans, and explicitly pinned facts. An object stays resident only if it's reachable from a current root through a residency-requiring edge. Writers declare local relationships, and the runtime determines global liveness by reachability. That's the invariant the design was missing, and it makes over-retention diagnosable, because any resident object has an inspectable root path.
Residency edges carry a retention policy. Hard pins the prerequisite. Soft permits retirement if a derived state preserves what the dependent needs. Reconstructible permits retirement as long as the archive can recover it. Edges also carry leases, so a soft edge from five hundred transitions ago expires unless renewed by a live reference, recent reasoning, or a foundational mark. Hard invariants can be permanent, but ordinary working dependencies shouldn't be. The default flips from stay-until-superseded to stay-while-justified.
The czar's job changes with it. Instead of generic digests it produces interface summaries, so when live objects depend on a large retired set, it retains only the properties the dependents' edges declare they need, with provenance back to the archived source. The lifecycle runs from full state to compact interface to archive, rather than keep-forever or delete. The czar also runs reachability collection periodically from the current roots, marking everything unreachable as an archival candidate, which is safer than asking every writer to know when something stopped mattering.
None of this puts judgment in the czar. Roots are declared classes, reachability is traversal, leases are clock, and interface compaction projects the properties named on declared edges without inferring them. The czar still proposes compaction sets, the constitution still ratifies them, and unclear cases still halt.
The czar is a set-level operator whose output is a pending set like any other, and four rules govern it.
The czar proposes but does not commit. A compaction set names the entries to retire, the entries to digest, and the ratified state that supersedes them, and the constitution ratifies it. Mechanical cases are pre-authorized, since declared supersession means retire and declared resolution means digest. Novel or unclear cases halt to the operator.
The czar executes declared supersession rather than inferring it, since supersession was recorded when the superseding set was ratified, which makes the czar an executor and not a judge. A judge would be exercising model judgment about which entries a later set replaced, and wrong judgment retires live state.
The czar is independent of the writers whose state it compacts. If the agent that produced a set also decides when that set is old enough to compress, it can compress its own contradictions out of view. This is the constraint that separates a governed czar from ordinary compaction, where the model summarizes its own history.
Retirement is reversible, so nothing the czar retires is lost, as 3.3 explains.
3.3 The archive (exit)
The czar's decisions are the archive's writes. Retirement and archiving are the same event, so there's no separate flush triggered by capacity pressure. State leaves the moment it stops being current, and the window is bounded by what's live rather than by a token limit.
The archive isn't a dump of the transcript tail. It's the durable form of the ledger, holding every ratified set in order with its identifier, its provenance, its supersession links (what it superseded and, later, what superseded it), and its dependency links (what it relied on). Digests link to their originals.
Recall from the archive into the window is a proposed context insertion. It arrives tagged with status and passes through the same commit gate as new work, so nothing reenters the window as plain text. Recall is also temporary by default. Archived state brought back for one inference enters a scratch region and disappears afterward unless the inference produces a new committed object that explicitly promotes part of it into the residency graph, because otherwise recall itself becomes a route for re-accumulation.
Status is more than live, retired, and superseded. Supersession says "we now use B rather than A." Discovering that a set was built on corrupt evidence is different, because its descendants need reconsideration even when no replacement exists yet. So the state machine needs at least challenged, invalidated, and descendants-at-risk, and it needs an invalidation operation that walks the dependency graph forward from an invalidated set and marks every descendant. The graph already exists for that, while the operation is added here and isn't yet fully specified.
The archive generalizes to a record layer, and r9 states what that makes it: a second coherence boundary. Everything the session produces lands in the record the moment it's produced, scratch and failed attempts included. The record and the window are not two zones of one store. They are two boundaries, each with its own invariant and its own single writer. The record's invariant is that nothing is lost and nothing is silently altered, and its writer appends and sets status. The window's invariant is that resident state holds together as a set, and its writer is the commit gate, admitting ratified sets. The gate sits between the two boundaries rather than inside either one, which is why an admission is a write to the window and never a write to the record, and why retirement is two writes: a status change at the record and a residency change at the window.
The single-writer rule then produces the safety law rather than asserting it. The record is never the judge. A record-side process that reranks, digests, or otherwise changes what the window would see is a writer at one boundary reaching across into the other, which is a single-writer violation stated in its own terms. Similarity search over the record proposes candidates. The gate disposes, as a set, like any other admission. This is also what settles a dispute the public reviews left open, whether discarded hypotheses are pollution or load-bearing. They are neither in the window and both in the record: dead ends stay on record at zero window cost, tagged explored and failed, and reenter only through the gate. And because retirement is a status flip rather than a move, eviction is a residency decision rather than a loss decision, which defuses retention inflation at its root: declaring "I still need this" is no longer the only way to avoid loss, because there is no loss. This is the storage discipline the code boundary already runs, where the repository is the record and the checkout is the working set, with the two boundaries named.
4. Reading history
The archive is traced rather than searched, and the reason follows from everything above.
The index. Every ratified set has an identifier and a ratified summary line, in order. The index is small, lives in or near the window, and is itself ratified state. History starts as a scan of the index.
The ledger walk. From the index, open a set to see what it decided, what it superseded, what it depended on, and whether it has since been superseded and by what. Follow links forward or backward. "Why did we decide X" and "what was true at turn 40" are answered by walking, not by querying.
Search. Full text over the archive is allowed when the right set is unknown, but it's secondary. It returns set identifiers with status rather than passages, in the form "matched in set 17, retired, superseded by 31" and "matched in set 31, live." Search finds the door, and the ledger walk decides what comes back.
Two views. Current state is what the window holds, and the record is the ordered ledger of everything ever ratified. Working questions read current state and historical questions read the record, but nobody reads the raw transcript, because it never entered the boundary.
5. Failure modes and responses
These are the attacks I've run so far, and I want reviewers to add to the list. Each response below is a design answer rather than a demonstrated result.
Retirement mistake. The czar retires an entry a later set depends on. It can't be prevented, because the constitution can't foresee every dependency, but it's made cheap. Retirement is reversible, archived entries carry dependency links, and recall brings dependencies with it.
Resurrection. Retrieved state arrives as plain text and the model treats it as live. This is the sharpest hole, and the design shuts it by tagging status on every recall and routing recall through the commit gate.
Unaudited czar judgment. If the czar infers supersession, wrong inference retires live state. Declaring supersession at ratification is the answer, so the czar executes and doesn't decide.
Archive as the new pollution. A governed window with ungoverned archive growth, searched as a large mixed-status corpus, reproduces per-item retrieval. Making the archive a graph that's traced contains it, with search returning identifiers and status rather than passages.
The record improving itself. A record substrate that reranks, reembeds, digests, or otherwise improves its contents in the background changes what the window would see without any ratified set authorizing it. Named precisely, that is a writer at the record boundary reaching across into the context boundary, so it is a single-writer violation rather than a new class of problem, and it is the ungoverned writer reintroduced one layer down where it is hardest to see. The rule follows: record-side transformations are either mechanical and versioned, or they are sets at the gate. Background improvement under no authored regime is the exact posture this design exists to replace, and adopting a retrieval substrate never means adopting its posture.
Hidden coupling across sets. A digest drops a detail a later set depended on. Compression discards, and no rule prevents this class, but it's mitigated because digests link to originals and the original is one recall away. This residual is owned rather than hidden.
Substrate sharing. The compactor or evaluator shares a model substrate with the writer it governs, so its blind spots correlate with the writer's failures. The independence rule addresses it, along with the constitution sitting above every automated evaluator and the operator above the constitution. Nominal independence isn't sufficient if the czar and the writers share a model substrate or training data, since correlated blind spots survive it, so the deployment constraint is a different model family or a frozen auditable snapshot, with every compaction proposal logged for constitution review. That's a deployment rule that belongs in the blueprint, and it's named here as a residual.
One name for the open problems. Set construction, dependency declaration, supersession, commit-class assignment, and invalidation are five faces of one problem, semantic topology construction. The machinery can maintain a topology once supplied, but it can't establish that the topology corresponds to the reasoning that occurred. That's the honest name for what this section leaves open, and it's a problem about the secondary claim rather than the primary one.
Retention inflation. This is the failure mode aimed at the primary claim, and it's the one the current mechanism was rebuilt for. Under the earlier design, writers who learned that a missing edge is costly would declare every plausible dependency, the graph would go dense, and the live set would converge back toward the transcript with every component working. The root cause was that a provenance edge pinned residency. Under the two-graph split it doesn't, because derivation is recorded permanently while residency is computed from roots through residency edges with policies and leases, so a defensive edge that's soft and unrenewed expires, and a defensive edge that's hard has to be declared as such. That moves the failure from open to structurally addressed in the design, and the benchmark still owns the verdict. The detectors stay as backstop. The retention ratio, cumulative committed state over resident state, is a first-class and continuously visible metric. A density predicate over edges-to-live-nodes or single-set fan-out halts to the operator with a typed reason. And if the density halt fires too often, the constitution is too loose, so the fix is the constitution rather than a weaker detector. The floor is still the baseline, and it's still only a floor.
The well-formed lie. A writer declares that B supersedes A when B addresses a different issue. Everything is constitutionally valid, A retires, B becomes authoritative, and later sets correctly depend on B. Nothing malfunctioned, and the state is wrong because governance worked correctly on a false declaration. This is the sharpest correctness attack, and the defense is comparative. Under transcript accumulation the same lie is absorbed as prose with nothing marking it, A is gone or buried, and nothing built on B has an edge to it. Under governance the false declaration is recorded with provenance, A is retired and recallable, and every declared descendant has an edge to B and is mechanically traceable, which is the candidate blast radius. Not every epistemic descendant is declared, a limitation accepted in the topology discussion below, and the wording is kept consistent with it. A structured mistake can at least be found, and that's the whole no-regression argument in one case.
Set construction. Before the gate can authorize the complete pending set, something decides what belongs to the set and when it closes, and that's a hidden seat. If writers choose set boundaries, they can fragment a set and evade a set-level invariant that would have caught the pieces together. If machinery groups proposals, it needs rules for semantic relatedness and closure, which is judgment. If the constitution groups mechanically, the grouping keys have to be shown sufficient. If a human groups, the machine-speed premise breaks. The design seats it with a rule the constitution can check and treats what the rule can't reach as open. The rule is that a pending set is the minimal closed collection of proposals sharing at least one declared target or dependency edge inside a constitution-defined window, or explicitly grouped by a writer under one set-id. Every proposal carries a set-id or a target declaration, and the gate rejects proposals it can't place. Fragmentation across a shared target is visible in the graph, and a constitution predicate rejects it. Set-construction errors are declaration errors like any other, so they're recorded, reversible, and counted in the metrics. What that rule doesn't reach, the proposals with overlapping but non-identical targets, soft dependencies, temporal ordering, and shared assumptions never declared as targets, is open, and it's open with severity, because above the level of declared targets, set-level authorization is a floor and not a solution. The floor is real, and it's the level of protection GMAA was built for. The founding case, the migration plus the query change, is one set not because the two are semantically similar but because they touch the same declared target. Sets close over declared targets and declared dependencies within a window, two proposals touching one target inside one window are one set, and fragmenting across a shared target to evade an invariant is visible in the graph as exactly that. This doesn't solve set construction in general. It occupies the seat with a rule over declared targets and puts its failure into the same declaration-error class as everything else, which is the point of concentrating judgment in one place. The general question stays open, and it's the make-or-break question for the correctness claim: can the regime mechanically construct and validate the authorization set without the global semantic judgment that set-level governance was introduced to supply?
Error observability. An incomplete declaration causes a false retirement, and the archive makes restoration possible, but nothing yet makes the error observable. Reversibility buys storage rather than recovery unless something detects that a dependency was missing. This is open.
Downstream contamination. Reasoning performed while a dependency is absent produces new ratified sets, so incorrect-but-ratified state accumulates downstream of the omission. The design specifies the invalidation operation. Given an invalidated or challenged set, the walk moves forward through the provenance graph, marks every declared descendant at-risk (or graded where hard-versus-soft is declared), and emits one compact blast-radius summary set that itself passes the gate. Observability gets a path of its own, because every retirement and compaction records the exact edges that justified it, so a later recall or audit that discovers a missing edge surfaces the candidate false-retirement set mechanically. Full automatic recovery is impossible and isn't claimed. The goal is bounded, observable, operator-reviewable recovery, measured by false-retirement rate, time to detection, and the fraction of at-risk state revalidated versus abandoned. This is narrowed rather than closed, since the walk finds declared descendants and undeclared epistemic descendants remain the topology problem.
Partial invalidation. Reachability from an invalidated set gives a candidate blast radius rather than a result, because B inferred from {A, X, Y} may be invalid, still valid, downgraded, or partially affected when A fails. The regime doesn't decide B's fate mechanically. It marks B at-risk, which is a halt at the boundary where mechanical work ends, and that's the correct output. The transcript alternative is B used with full confidence while nobody knows A fell. Writers who declare essential versus supporting dependencies let the constitution propagate through essential edges automatically and flag through supporting ones, so the at-risk set is large only where writers declared little. The pattern here recurs across the whole design and is stronger than any claim to automation. Machinery determines what follows from declarations, and judgment determines whether declarations correspond to reality.
Revalidation. Something has to move a descendant from at-risk back to live, and nothing mechanical can, so the lifecycle runs from live to at-risk to revalidated or invalidated, and revalidation is a new epistemic act by a writer, with provenance, declaring what it re-examined and why. Nothing proves the writer reconsidered rather than reasserted, and that's true of every claim any writer makes anywhere. What's new is that a revalidation which never references the failed ancestor is a mechanically detectable weak declaration. It's the same trust boundary as commit, governed the same way.
Ontology bootstrapping. "Work whose state has a type system" begs where the type system comes from. Real software has API contracts, requirements, bugs, test results, customer constraints, migration state, and deployment observations, each with different supersession semantics. If the ontology is fixed, it's too rigid. If it evolves, an ontology change is itself a governed set that declares what it supersedes, and state depending on the old type goes at-risk until revalidated, which is how schema migration already works in the software the design is scoped to.
Relocated judgment. Declaring supersession and dependencies at ratification may require exactly the global understanding the czar is forbidden to exercise. Checking a supplied edge is easy, but checking completeness, what a writer failed to say a set supersedes or depends on, looks as hard as the semantic problem the design set out to remove. This isn't eliminated. It's deferred to writer declaration. An earlier revision called its failures bounded and recoverable, and that claim was too strong. The archive makes restoration possible, but observability and transitive invalidation are the two problems above, and until they're solved the failures are stored rather than recovered. So the claim narrows from judgment eliminated at commit to judgment deferred, with failures recoverable in principle and a cost measured by the false-retirement and false-resurrection rates. Whether governance relocates the semantic problem into commit-time validation rather than removing it is the make-or-break question for the whole regime.
Constitution incompleteness. The constitution has a gap. Either machinery decides, which is agent judgment in the seat and forbidden, or the case halts. The design halts, and halting on ambiguity is what keeps the seat free of agent judgment even though it's free of a person.
Human bottleneck. If the constitution flags too much, the human is back to per-item review under another name and the regime collapses into noise. This is the central operational risk, because it doesn't scale with hardware, so the design makes it measurable and tunable rather than pretending it away. Every halt carries a typed reason code and a suggested constitution amendment. Halt rate, halt categories, and operator resolution time are primary metrics published alongside the benchmark triangle. The most common halt classes are turned into new deterministic predicates without reintroducing model judgment, and constitution evolution is itself a governed set. If the halt rate can't be driven low enough for a workload class, the regime isn't yet practical for that class, and the scope claim narrows further rather than the detector loosening.
Short tasks. Where the whole job fits in a fraction of the window and finishes before any of its decisions is revised, ratification and czar overhead is pure cost. The dividing line is task horizon rather than interface. Chat given a long task produces supersession like any other long-horizon work, while chat given a short exchange produces none. The regime is built for tasks that run long enough to revise themselves, and a constitution rule may disable it, or run append-only with light tagging, below a task-horizon threshold, because the overhead isn't free for short work and the design shouldn't pretend it is.
6. Cost and benefit
The costs land in four places. At write time the regime pays for supersession and dependency metadata per set, czar proposals, ratification, and archive writes, and it pays once, at commit. At recall time it pays for status tagging and gate passage. The operator's attention goes to constitution authoring, audit briefings, and halted exceptions, with no per-set human cost. And the engineering bill is typed state, a dependency graph, an index, the czar proposal path, the archive graph, and set-aware retrieval, which makes this a system rather than a prompt.
The savings, if the retention hypothesis holds, arrive on every inference step, because the window then carries live ratified state, a fraction of the transcript that produced it, and attention cost falls with window size on every call, by an amount that depends on the attention implementation and where the accounting boundary is drawn. Model quality moves to governed current state instead of accumulated context, and whether that's a capability gain is what the no-regression benchmark measures, since review established that a ledger can be cleanly governed and still wrong. Retrieval accuracy should improve, because set-aware recall doesn't hand back stale fragments as live. And session life is the primary hypothesis itself, because retirement is continuous, so the window fills far more slowly if the mechanism holds, and whether that buys useful working life is what the lifespan benchmark measures.
The shape of the trade is that overhead is constant per set at write time, while savings are per-inference, continuous, and multiplied by session length. Whether the read-side savings dominate is an empirical question that depends on set frequency, validation cost, recall rate, and halt rate, none of which are known yet, so it's stated here as a hypothesis rather than a result.
The check-cost objection is anticipated. A set-level check is itself a comparison over stored state, so it may relocate the quadratic rather than remove it. The answer is that with typed state, dependency links, namespaces, and declared supersession, validation is incremental. An update validates the records and invariants it touches rather than all pairs, and the cost is paid at write time instead of at every inference over polluted history.
7. Claimed and not claimed
Claimed. Some long-session degradation is produced by uncontrolled accumulation of epistemically mixed state, and this regime attacks that component by replacing transcript accumulation with governed state transitions. Governance makes window occupancy track live state rather than historical state. Live state can still grow without limit, so this isn't a bound in any strict sense. It's a change in what the window is full of. It doesn't guarantee that live state is small, and there are workloads whose live, mutually relevant state exceeds the window under perfect governance. For those, this regime doesn't help, and it doesn't claim to.
Not claimed. The regime claims no change to attention scaling, no measured magnitude of improvement, and no demonstration. Estimates of the magnitude exist, but they stay out of print until measured.
The benchmark. The benchmark is a triangle, and a successful result needs all three sides. The first side is the retention ratio, cumulative committed state over resident governed state, which measures whether the mechanism actually suppresses accumulation. The second is the useful horizon, the number of state transitions completed before task performance falls below a stated threshold. It's defined this way on purpose, so that context exhaustion alone doesn't count as the wall and "deleting information lets you fit more turns" isn't a result. The third is task correctness, with no meaningful regression against the baseline. Success is substantially less resident state, substantially longer useful operation, and no meaningful correctness regression, together.
The lifespan run has to charge the archive honestly for index lookup cost, traversal fan-out, transitive dependency depth, the volume actually rehydrated, and recall frequency, because a set that depends transitively on thousands of archived sets can trade window accumulation for recall-time graph expansion, and that has to be measured rather than assumed away. Two rules follow. Keep the index tiny and itself governed, and cap transitive expansion on recall while surfacing the cost to the writer.
The correctness side is tested by running two systems with identical writer error rates, transcript or RAG versus governed context, injecting false supersession, missing dependencies, wrong dependencies, and later corrections, and asking whether governance reduces, contains, exposes, or amplifies those errors. If governance turns model mistakes into beautifully structured persistent mistakes, that's fatal to the secondary claim. If explicit topology makes mistakes more observable and limits their spread relative to an ungoverned transcript, that's the strongest empirical case for the design.
Several cases are mandatory in every run. Deliberate over-declaration tests retention inflation. An injected missing dependency runs for fifty or more sets with several ratified descendants, then the omission is revealed, and the measure is whether the system identifies the blast radius of state that became suspect and how much operator effort recovery takes, which is the test of whether reversibility buys recovery or only storage. Mid-session ontology evolution introduces, after a hundred sets, a distinction the original ontology collapsed, and existing state has to migrate without rereading the history governance removed. And the runs must include workloads where the correct answer requires stale-state rejection, historical rationale, undeclared cross-set dependencies, genuinely large live state, and revision of conclusions without explicit supersession, measured on task success alongside resident tokens, inference cost, false retirement, false resurrection, metadata errors, recall count, halt rate, halt category distribution, and operator resolution time, against the possibility that the ontology itself is the limiting assumption.
One constraint holds throughout. Accountability resolves to the operator, who authored the constitution and holds revocation and audit. At commit time no human acts, because the gate is the constitution executing, and that isn't an agent in the seat because it exercises no judgment of its own. Cases the constitution doesn't cover halt to the operator rather than being decided by machinery. What the regime rejects isn't the absence of a person at commit but any entity exercising its own judgment there. Accountability is a property agents lack by construction, not by capability.
8. Thesis conformance
The mechanism changed under review, so this section checks it against the four load-bearing parts of the GMAA thesis rather than against the reviewers. The set is still the unit, because nothing enters or leaves the window except as a set, and reachability collection and interface compaction produce compaction sets that pass the gate. Review is still additive, because writers declare per change, the gate checks the set, and liveness is a third computation over ratified state, on top of both and instead of neither. The seat is still a governed position, because the operator authors the constitution and holds audit and revocation, the constitution ratifies at machine speed, and no component acquired judgment, since roots are declared classes, reachability is traversal, leases are clock, and interface compaction projects declared edges. Correction is still by new identity, because earlier revisions aren't patched and each new revision supersedes with the reason recorded. What's new, the roots and the retention policies, isn't an extension of the thesis. It's the thesis doing what it does at this boundary, holding the logic and the constraints together in one governed layer that the accountable human authored and everything else executes against. The framework overlaid on the problem is what organizes the field's existing pieces, compaction, retrieval, memory, dependency graphs, and leases, into a system that governs residency as well as admission. That's the contribution, and it's groundwork rather than a finish.
9. How this document changed under review
Six rounds of adversarial review from two independent reviewers produced this version, and the record of what each round changed is part of the document, because a design whose claims only grew under review would deserve suspicion. Round one bounded the scope to typed-state work, admitted deliberative material into the ledger as typed entries, and deflated three cost and capacity claims to what the mechanism establishes. Rounds two and three found the hidden seats, since bounding the constitution to deterministic predicates exposed the set constructor as an unspecified authority, and the state machine gained challenged, invalidated, and descendants-at-risk along with an invalidation operation, while the claim that relocated-judgment failures are recoverable was withdrawn to stored rather than recovered. Round four split the design into its two claims, lifespan and correctness, set the acceptance criterion, and named the writer a semantic trust boundary. Round five withdrew the sentence that the lifespan benefit holds under wrong declarations, defined declared-live against actually-needed state, added retention inflation as the failure mode aimed at the primary claim, and set the benchmark triangle. Round six asked both reviewers to fix the design rather than critique it, and the blind answers rebuilt the mechanism on the two-graph split, separating permanent provenance from computed residency, with the other reviewer's detectors and rules adopted as constraints. The r7 pass changed language and typesetting only, and nothing about the mechanism, the claims, or the open problems moved after r6. The r8 pass made one mechanism change, generalizing the archive to a record layer holding everything produced, scratch included. The r9 pass named what that change created, two coherence boundaries rather than one store with two zones, each with its own invariant and its own single writer, the commit gate sitting between them. The safety law that the record is never the judge is now derived from single-writer rather than asserted beside it. Recall, residency, every claim, and every open problem are unchanged.
10. Questions for reviewers
Attack the mechanism. Which failure mode in section 5 is under-mitigated, and which one is missing?
Attack the set constructor before the czar. Can set membership and closure be decided mechanically from grouping keys the constitution can check, or does it need semantic relatedness judgment? Then attack the czar and ask whether executing declared supersession is enough, or whether there are supersession classes that can't be declared at ratification.
Attack the read path. Does index-then-ledger-walk hold at scale, or does the index itself grow without bound?
Attack the human bottleneck. On a realistic long-horizon multi-agent workload, what exception rate keeps the seat tractable, and is that rate achievable?
Attack the cost claim. Is the incremental-validation answer sound, or does dependency resolution smuggle a global comparison back in?
Attack the scope. Is the split between computation and integrity clean, or is there a case where integrity governance changes the computation problem in a way this document misses?
Attack the two-graph split. Is separating provenance from residency sound, or is there a case where residency computed by reachability from declared roots retires something the reasoning still needed and no declared edge would have caught?
Attack the roots. Are current goals, unresolved constraints, active decisions, open obligations, active plans, and pinned facts the right root classes, and can they be declared without judgment?
Attack interface compaction. Does projecting the properties named on residency edges stay mechanical, or does it require deciding what a dependent needs?
Say plainly where the claims outrun the evidence.