How GMAA Works
Governed Multi-Agent Architecture: the process, for anyone deciding whether to build on it. Canonical: gmaa.ai · Spec licensed CC BY 4.0 · cite by version
Everyone built the pipes. No one built the chain of command.
Point several capable AI agents at one codebase and the failure isn't necessarily any single agent being wrong. It's coherence collapse: two changes, each correct on its own, break something together. A migration and a query that each pass their own tests but disagree about a column. Work that's individually green and jointly broken. Nobody was accountable for the combination, because everybody reviewed their own piece.
Most frameworks give you better pipes, faster agents, richer tools, more orchestration. GMAA gives you the thing that was missing: a chain of command over a shared system of record, so that many agents can propose, but the set of what they propose is judged as a whole before any of it lands.
The idea: authorize the set, not only the individual change
GMAA's one claim is that multi-agent systems fail by authorizing at the wrong unit. They approve changes one at a time, while the coherence of the complete pending set goes unexamined.
So GMAA changes the unit. The thing that gets authorized is the set, the complete collection of pending changes, evaluated together against the system's rules, and ratified by one accountable human who is independent of every agent that proposed the work, before any member commits. Over each boundary of state that must stay coherent, exactly one seat commits; everyone else reads and proposes.
Two consequences follow, and they're the spine of the whole design:
- Authorization is not a probability. A confidence score is not a decision. The correct resolution of a jointly incoherent set often turns on a fact that lives outside the system (a planned cutover, a hold, a commitment to a customer) that no automated evaluator has a channel to read. So the final say belongs to a human, at a vantage point outside the boundaries being changed.
- Nobody signs their own work. The seat that builds a change never verifies it; the seat that assembles a set never ratifies it. Judge and judged are never the same context. That separation is enforced by structure, not by good intentions.
The cast
- The architect, where a human's intent enters. Works with the operator to decide what to build and hands that decision to the coordinator. It can live in a chat window or graduate into the codebase itself.
- The operator, the accountable human. Authorizes the work up front and gives the final go on every set.
- Foundation, the coordinator: the one seat that commits to the shared spine, and the hub every message routes through. Every other seat talks to foundation, never to another seat.
- The executor, the working seat that runs the side-effecting operations a change needs: migrations, live probes, calls out to databases and APIs, each inside a pre-authorized allow-list. Foundation authors what runs; the executor runs it, emits a receipt for every change to real state, and surfaces the result. It works on its own branch and never commits the spine, and it never widens its own credentials.
- The lanes, the working seats the architect and operator define at instantiation, one per territory of work that needs its own writer, plus any independent tracks they choose to run in parallel. Each is isolated in its own branch, proposes, and never commits the spine. They follow boundaries, not workload: a project gets exactly the lanes its territories call for, and more work alone is absorbed inside a seat, not by adding one.
- The auditor, an independent seat that reproduces every claim of "done" rather than trusting it.
How a project begins
Before any work is handed out, the architect and operator author the project's foundation: the specification, the contracts, the roster of seats, and the invariants that boundary has to keep. Foundation commits that foundation to the shared spine first. Only then does it distribute the first tasks to the working seats.
The order is the point. Every seat boots against a committed record rather than a conversation, so there is one source of truth from the first commit forward, and no seat ever works against intent that was never written down.
How a change moves
- A decision is made and handed to foundation. The architect and operator decide; the decision goes to the coordinator, not to a worker.
- Foundation records the dispatch, routes it to the seat that owns the affected area, and signals it to start. That owner is a lane where the project separates work into territories, or the executor where the change touches real state. The dispatch lives in the owning seat's inbox on the record, not in a conversation.
- The owning seat works in isolation, on its own branch, never touching the shared spine, never able to disturb another seat.
- It produces a result and a receipt, a record of what it did and proof its checks ran.
- Its checks run. Pass, and the change is verified. Fail, and it loops back for rework, a normal, expected part of the flow, not an error state.
- An independent auditor reproduces the check. Nobody signs their own work: a change is not "done" on the worker's own say-so, only when a different seat re-runs the check and gets the same result. The auditor reports that verdict; it never merges or closes the change itself. There is, by construction, no shortcut from "in progress" to "closed." And clearing this gate is not the same as landing: an audited change still waits for its whole set to be ratified before it commits.
When a seat hits a conflict it can't resolve, it doesn't guess. It records the conflict and surfaces it to foundation, which re-checks it: was this a real conflict, or did the seat make a mistake? A mechanism question, foundation resolves on its own. A question about the rules, what the contract should say, goes up to the architect and operator. Foundation decides how; the humans decide what.
The set and the gate
-
Foundation assembles the verified changes into a set and reasons over it. This is the heart of the design, not a rubber stamp: - It runs a conflict sweep across the whole set and the checks that are only meaningful on the combination (the kind of conflict that no single change reveals). - It classifies what it finds as in-band (settleable from facts already in the record) or out-of-band (needs a human fact, a cutover, a hold, a commitment). - Crucially, the rules it reasons against aren't asserted; they're probed. Each invariant has a live check that must actually fail on a bad input, not just pass on a good one. An invariant's health is its test streak, not its prose. This is the piece most governance frameworks lack.
-
The operator decides. - If everything checks out, foundation asks a simple question: "all checks pass on the set, commit?" A clean assessment only means nothing in the record objects, so this is a real gate, not a formality. The operator can still say no on a fact the system couldn't see. That "wait, what are you doing?" is exactly why the human is here. - If there's a conflict, foundation presents it with grounded alternatives, and the operator makes a real decision: ratify, ratify in a specific order, amend, hold, or reject. An amended set is a new set: it's re-assembled and re-checked as a whole. There is no line-item approval; that would quietly reintroduce the very per-change authorization the whole design rejects.
-
Foundation commits and makes it durable. On the operator's go, the one writer merges the set and pushes it so "landed" means truly landed, not just local.
What the assessment reasons against
The rules aren't a single file. A set can be found in conflict with any layer of a real governing body that ships with the engine:
- The law, the architecture and its universal invariants (live-data-only, one writer per boundary, one role per session, no fake closures).
- The contracts, how changes are authorized, how writes are mediated, how work is signalled and retired.
- The database itself, the schema is a governed, contracted surface, and the only way agents write to it is through stored procedures. The identity the agents run under is read-only by default, so it cannot directly insert, update, or delete; every write goes through an EXECUTE-granted procedure. A direct write fails by design. And this isn't a promise on paper: a standing probe continuously confirms the workload identity still cannot write around the mediation. The tables, columns, procedure signatures, and grant model are each a contract an adopter's changes are checked against.
- The schemas, the lifecycle a change must traverse, and the shape of a valid receipt.
- The disciplines, how a seat is expected to behave: verify before acting, halt loud rather than guess, prefer prior art.
- The enforcement, the auditor and the standing probes that check every invariant continuously, in the code and in the database.
- The project's own rules, the domain invariants each adopter declares when they stand up their fleet.
The reasoning is done by foundation, over evidence from the probes and the sweep, against all of that, with the human on the set as the final layer.
Why it's built this way
- Prevention over detection. "The harness would catch it" is only ever in addition to a structural partition, never instead of one. GMAA prevents the confused decision from happening, rather than hoping to catch it after.
- Structure over vigilance. One writer per boundary, identity that comes from where a seat runs rather than what it's told, an independent auditor by construction: these hold whether or not anyone is watching.
- The human where machines can't reach. Automated checks settle everything in the record. The human settles what isn't in it. Neither is asked to do the other's job.
Free and paid
The architecture is runtime-independent: the engine ships as two lines, one for Claude Code and one for Grok Build, both pinned to the same canon version and carrying the same gates. Which runtime you use does not change anything on this page.
GMAA is open-core. The specification is free and citeable: a standard you can adopt and build on. The engine, the runnable blank template, is free to use, modify, and redistribute, including inside your own commercial operations. It is not open source in the usual sense: you may not sell it or offer it as a paid or hosted service. What is paid is the genuinely hard part: turning that template into a coherent, running fleet in your real environment, the graduation from a template to a governed system that actually holds.
Read the spec, run the engine, cite it by version at gmaa.ai. For how the engine implements the specification, file by file with the execution wiring, see the engine map.
GMAA · Governed Multi-Agent Architecture · a chain of command for AI agents on a shared system of record.