GMAA Governed Multi-Agent Architecture
From the field-evidence log

What it has caught

Incidents the architecture caught while running in production, the subset cleared for publication. The evidence for the set comes first. A separate section records the governance holding on its own authors, which is a different claim.

A note on what is shown. GMAA keeps a running field-evidence log of incidents the architecture has caught in operation. Most of those entries stay internal, because they carry project-specific or proprietary detail that cannot be published: the client, the domain, the implementation. What appears below is the subset cleared for publication, rendered as pattern only: the shape of the failure and the shape of the catch, with the operational specifics held back. It is a sample of the record, not the whole of it, and it is offered as existence, not as proof of completeness.

Two changes passed their own reviews and still broke each other. Reviewing them as a set caught it before either committed.

The first time the complete pending set was assembled and judged as a whole, it surfaced two changes that had each been approved on their own, and whose authors had both signed off from memory, as jointly incoherent. It was caught before either one committed, at no cost to the system. Per-change review had passed both; only judging them together as a set found the break.

The same failure was already running in the system, unnoticed.

The failure the set gate prevents was not hypothetical. A set of pending changes had been accumulating with every member reviewed on its own and none of them evaluated together. Per-change review ran correctly at every seat; the set was the one thing nobody was scoped to check. It was the exact gap the argument describes, running unmeasured because nothing was instrumented to see it. Naming it is what turned an invisible risk into a governed one.

Three correct changes combined into a broken result, caught before any of them committed.

Three mechanisms, each correct on its own and each independently approved, combined into a state none of them intended: a stranded batch that no single change could account for. The break lived only in the set. It surfaced at the commit boundary, was assembled and judged as a whole under independent authority, and was resolved at the set level rather than patched at any one member. This is the core claim recurring in production: the unit that needed authorizing was the interacting set, not any single change.

Other builds reinvented set-level review on their own.

Separately, independent product builds arrived at the same set-level approval surface on their own, with no mandate to, because approving one change at a time kept producing the exact defect the argument predicts: a rubber-stamp surface and an unauditable batch. One of them stated the lesson plainly: a surface that can only say yes is a rubber stamp by construction. The label was deferred; the mechanism arrived anyway.

Checking each change on its own worked, and still was not enough.

In the same window, the per-change checks did their job: a fabricated input was refused before it reached live data, a stale local copy of a tool was caught, and a defect was stopped by a probe before any live run. Those checks are necessary. The set-level failure above, from the same window, is why they are not sufficient. Both gates fire, and neither replaces the other.

What stays open

These are classes the boundary has caught, offered as existence, not as proof of completeness. Whether any given build catches every joint incoherence is a separate, open question with a named way to earn the answer: an adversarial test that hunts for the miss. It is called complete when that test comes back clean, and not before.

The governance also polices its own authors

The three below are a different axis. They are not evidence of the set. They are evidence that the governance holds even on the seats that write its rules, which is a claim about independence, not about the unit of authorization. They sit in their own section on purpose, because a claim about the set should be carried only by the set.

The system caught its own rule-author.

After a context restart, a seat refused to act on an instruction because the rule it cited existed only in chat and had never been committed to the record. Nothing at the system of record backed it, so the seat halted and surfaced the gap rather than proceeding. The mechanism held its own governance traffic to the same evidence bar as worker output.

The evidence record caught its own entry.

An append to the field-evidence log itself was composed from a stale local copy and arrived out of sequence. The pre-commit check against the live record caught the collision, held the write with zero corruption, and corrected the sequence without altering the author's text. The discipline the log documents caught the log's own append.

The record corrected the reviewer.

A ruling cited a mechanism as the fix. A live read of the actual system showed that mechanism could not do what the ruling assumed. The layer corrected the mechanism to match the ruling's intent, under independent authority, without relitigating the decision. Substrate beat artifact, and it bound the rule-author too.

Field evidence, cleared subset · offered as existence, not completeness · cite by version