Orchestration Engine — Conceptual Deep Dive
Purpose & Mental Model
Agentweaver orchestration answers one question: how does a high-level goal become safe, reviewable, mergeable work performed by a team of agents?
The engine is intentionally split into two layers:
- Coordinator orchestration decides what should happen. It turns an ambiguous goal or backlog item into a confirmed outcome, decomposes that outcome into a dependency-aware plan, assigns work to team members, and assembles the results.
- Run workflow orchestration decides how each run moves through gates. It applies a declarative workflow to live execution: agent work, safety review, human review, merge, and scribe recording.
The split matters. Planning and decomposition need durable state, idempotency, and team-level reasoning. Individual run execution needs streaming, review gates, restart loops, and terminal status handling. Keeping those concerns separate lets Agentweaver recover from partial progress without re-asking the model to re-invent the plan.
A useful rebuilding rule is: the coordinator owns intent and coordination; workflows own execution gates.
Casting and Blueprints feed orchestration with team shape, role charters, and workflow defaults. Blueprint validation accepts only review_policy: default; this is not a configurable project review-policy overlay. They are summarized here only; the detailed explanation lives in team-casting.md.
Live workflow execution and non-workflow single-prompt paths use the Copilot SDK. The live workflow worker path does not switch worker implementation based on the run's model source; custom providers are configured through the SDK.
Core Design Invariants
These invariants are the backbone of the system:
- Persist decisions before doing work. The requested outcome and work plan are stored before child runs are launched. Recovery starts from persisted intent, not from chat history.
- Confirm ambiguity at the boundary. The coordinator may draft, revise, and ask for confirmation before committing a plan. Once confirmed, later components can assume the outcome is intentional.
- Use declarative graphs for policy. Workflows describe nodes, gates, and edges. Runtime code binds those declarations to executable steps and fails closed when a step cannot be safely bound.
- Advance only the ready frontier. Subtasks form a DAG. A subtask can run only after its dependencies are complete, so parallelism is safe and deterministic.
- Separate child work from collective responsibility. Child runs produce pieces ready for assembly, not independently reviewed pieces. The parent coordinator assembles, reviews, merges, and records the combined outcome.
- Make gates explicit and durable. Safety, human review, merge, and terminal states are visible run states and stream events, not hidden control flow.
- Prefer idempotent recovery over clever replay. If a plan already exists, reuse it. If a run already reached a gate, resume from that gate. If a stream disconnects, replay durable events.
Coordinator Orchestration
Problem It Solves
A user often gives Agentweaver a goal, not a task list. The coordinator converts that goal into something a team can execute safely:
- What exact outcome are we trying to produce?
- What assumptions or constraints define success?
- Which parts can run independently?
- Which specialist should own each part?
- What must be reviewed before changes merge?
Without this layer, every agent run would independently interpret the same broad request. That leads to duplicated work, conflicting edits, and unclear ownership.
OutcomeSpec: The Intent Contract
The first durable artifact is the OutcomeSpec. Conceptually, it is the contract between the requester and the system.
It captures:
- the original goal,
- the desired outcome,
- scope and exclusions,
- assumptions,
- clarifying questions or revision feedback,
- and whether the outcome is still being drafted, awaiting confirmation, confirmed, or declined.
The important design choice is that confirmation happens before decomposition is treated as authoritative. The coordinator can draft an interpretation, receive revision feedback, and loop until the requester or unattended policy confirms it.
Because the confirmed OutcomeSpec is the source of truth for intent, it — not the later decomposition — is where the work's breadth originates. When the drafter frames the outcome it reads the project's team roster and honors the breadth the goal explicitly asks for: a full-lifecycle "from the initial idea through to a working app" goal yields an outcome whose scope enumerates the discovery, PM, design, and build deliverables the goal warrants (filtered to what the team can actually produce), while a narrow goal stays lean. Downstream workflow selection and decomposition then faithfully honor that confirmed breadth, so if the outcome is narrowed at drafting time, the whole PM/discovery half is silently dropped everywhere after it. See How drafting works for the roster-awareness and scope-breadth guidance that shapes the draft.
Rebuild guidance: treat the OutcomeSpec as the source of truth for intent. Do not let individual worker agents reinterpret the original request independently once the spec is confirmed.
WorkPlan: The Execution Contract
After confirmation, the coordinator creates a WorkPlan. The WorkPlan is the execution contract for the parent coordinator run.
It stores:
- the confirmed OutcomeSpec it implements,
- the selected workflow,
- subtask records,
- dependency edges between subtasks,
- assembly state,
- and any integration branch or coordination metadata.
Each subtask includes its assigned agent, model choice, charter/context, isolation intent, status, child run id, and any recovery guidance.
The plan is a DAG because ordering is a correctness constraint. If subtask B depends on subtask A, B should not start merely because an agent is free. This allows safe parallelism: every tick can dispatch all currently-ready nodes while preserving required sequencing.
Rebuild guidance: store the plan before dispatch. If the coordinator crashes after planning but before child runs start, it should resume from the persisted WorkPlan rather than ask a model to decompose again.
Coordinator Control Flow
The coordinator flow has two phases:
- Model-assisted planning phase — draft and confirm the OutcomeSpec, select a workflow, decompose the work, and persist the WorkPlan.
- Service-driven execution phase — dispatch ready subtasks, watch child runs, assemble results, and advance the parent run through review and merge gates.
The coordinator is designed to be idempotent. If it is asked to orchestrate a run that already has a WorkPlan, it does not create a second plan. That invariant prevents duplicate child runs and conflicting DAGs.
Decomposition Logic
A good decomposition algorithm should produce subtasks that are:
- owned by one agent,
- bounded enough to complete independently,
- ordered by explicit dependencies,
- labeled with intended isolation or file ownership,
- and recoverable with enough guidance to retry or inspect failures.
Agentweaver treats dependency edges and file/isolation hints as coordination data. The dependency graph is the hard ordering rule. Isolation hints are advisory: they help avoid conflicts and guide dispatch, but they are not a substitute for merge conflict handling or review.
Cycle breaking is essential. Model-generated plans can accidentally create circular dependencies. A production coordinator should detect cycles and either remove weak edges, ask for clarification, or fail before dispatch. Dispatching a cyclic plan would deadlock because no frontier can become ready.
Dispatch and Assembly
The dispatcher repeatedly asks: which pending subtasks have all dependencies completed? Those subtasks form the ready frontier.
For each ready subtask, it launches a child run with an isolated working tree and output branch. Child runs are intentionally trimmed: they perform agent work, then stop at an assemble-ready boundary. They do not each perform RAI, human review, merge, or scribe. Those are parent-level responsibilities because the user reviews the combined outcome, not a pile of isolated fragments.
When a child reaches assemble-ready/completed, the dispatcher rebuilds the coordinator integration branch from the successful child branches in dependency order. Dependents are then branched from that integration branch, so they can read files produced by their prerequisites without concurrent siblings sharing one mutable git index.
Assembly is where the coordinator turns independent child outputs into one coherent result. This is also where conflicts, missing pieces, and cross-subtask inconsistencies should be detected before the parent enters review and merge gates.
The child graph branches on AgentTurnOutput.TerminalFailureReason: a clean turn reaches child-assemble-ready, while a typed failure reaches child-turn-failed (apps/Agentweaver.Api/Runs/RunWorkflowFactory.cs:788–814). “All subtasks settled” is not equivalent to “all outputs eligible for assembly”; a failed child is not an approved aggregate input.
Where this lives:
apps/Agentweaver.Api/Coordinator/apps/Agentweaver.Api/Memory/
Workflows and Trigger Evaluation
Workflow as Policy Graph
A workflow is not just a list of functions. It is a policy graph that describes how a run should progress through work, checks, review, merge, and terminal states.
A workflow definition answers:
- What starts the graph?
- Which node performs agent work?
- Which gates can send work back for revision?
- Which failures are terminal?
- Which path means success?
- Which event or schedule declarations can initiate backlog work for this workflow?
The shared illustration describes the standalone built-in workflow, not mandatory collective assembly policy. Its current success path is agent -> rai -> review -> merge -> push-pr -> scribe -> done (apps/Agentweaver.Api/Workflows/DefaultWorkflowTemplate.cs:42–161); child runs bypass this graph.
The important idea is that loops are first-class. Safety or review can return work to the producer. Merge can return to review if blocked. Terminal failures are explicit exits, not exceptions swallowed by the runtime.
Invocation context and event triggers
Manual and heartbeat origins are recorded as invocation context. They do not remove valid workflows from the selector candidate set. Event and schedule triggers are evaluated before the normal backlog and coordinator pickup path. A requested override is used only when it resolves to a valid, bindable workflow.
Workflow Selection Logic
The selection order is deliberately conservative:
- Load built-in, catalog/library, and project-authored workflows.
- Record invocation kind from the run origin.
- Honor a valid override if present.
- Order the configured project default first without short-circuiting automatic selection.
- If exactly one workflow remains, use it without model help.
- If several remain, ask the selector to choose the best process fit.
- If selector output is invalid or parsing fails, fall back safely rather than inventing a workflow id.
This pattern limits model authority. The model may choose among safe candidates, but it cannot bypass validation or runtime binding.
See workflow selection for override events, selector retry/fallback behavior, and the post-decomposition Build & Test compatibility check. That check can choose a platform software workflow when automatic project candidates cannot cover code-producing work; an explicit workflow lacking Build & Test is honored with a warning.
Binding Declarative Nodes to Runtime Execution
A workflow file describes intent. The runtime must bind that intent to concrete executors.
The binder should:
- classify nodes by type and gate kind,
- resolve each node to a known executor,
- expand logical edges into the live execution graph,
- verify every workflow-declared gate and transition has a binding,
- and fail closed if a required node cannot be executed safely.
Failing closed is a security and correctness property. A workflow that asks for a safety gate but cannot bind one should not silently skip safety. Likewise, a custom node type should not become a no-op merely because the binder does not understand it.
Some workflow shapes, such as fan-out/fan-in style nodes, are design-level extension points: the graph model can express them, and the binder is the place where their executors are resolved. They give the system room to grow more complex execution patterns without changing the surrounding contract.
Where this lives:
apps/Agentweaver.Api/Workflows/docs/workflow-library.mddocs/workflow-binder.md
Run Lifecycle
What a Run Represents
A run is the durable unit of execution. Conceptually it bundles:
- a project and workspace/worktree,
- the assigned agent and charter/context,
- the selected workflow,
- the run origin,
- live and durable event streams,
- and a persisted status.
A run can be started directly by a user, reserved by backlog pickup, created as a coordinator parent, or launched as a coordinator child. All forms should converge on the same lifecycle machinery so status, streaming, review, and recovery behave consistently.
Parent, Child, and Pickup Runs
Agentweaver uses run origin to preserve intent:
- Manual runs are user-started and usually go through the full workflow.
- Coordinator parent runs own the team-level plan, assembly, review, merge, and scribe phases.
- Coordinator child runs execute one subtask and stop at the assemble-ready boundary after agent work.
- Backlog pickup runs are coordinator runs created by the heartbeat loop for unattended ready tasks.
The key difference is not the storage shape; it is the responsibility boundary. Child runs should not merge independently because they are fragments of the parent outcome. Parent runs should not redo child work because they coordinate, assemble, and gate the whole result.
State Machine
This is a conceptual, non-exhaustive state machine. It emphasizes externally visible gates. When a human review node is reached, the run becomes AwaitingReview and the client can act. When merge is requested, the run becomes Merging. These are not merely internal events; they are durable states used by clients, recovery, and monitoring.
Runtime Sequence
The watch loop translates live runtime events into persisted run state. This keeps state transitions centralized. The agent produces work; the workflow emits events; the watch loop decides what those events mean for durable status and client-visible stream completion.
Event Streaming
Run events have two purposes:
- Live feedback — clients can see what the agent is doing now.
- Recovery and reconnect — clients can replay what happened if they disconnect or the process restarts.
The Postgres path appends durably, then reads ordered rows after the subscriber's cursor:
EfRunEventStream allocates the next sequence under a per-run advisory transaction lock and acknowledges only after commit. Subscribers query Sequence > cursor and poll again after 250 ms when no rows are available. A reconnect can therefore land on another API replica without relying on the first replica's channel. The SQLite/local alternative has a bounded process-local channel; that channel is not the Postgres cross-replica delivery mechanism. The SSE endpoint supplies framing and completion behavior.
Where this lives:
apps/Agentweaver.Api/Runs/packages/Agentweaver.AgentRuntime/Workflow/apps/Agentweaver.Api/Infrastructure/docs/run-event-stream.md
Backlog and Heartbeat Pickup
Problem It Solves
The backlog lets Agentweaver accept work before an agent is actively assigned. The heartbeat loop turns ready backlog items into unattended coordinator runs.
This separates commitment from execution:
- A task can be captured and ordered in the backlog.
- Later, when it becomes ready and workspace conditions allow, the system claims it.
- Claiming creates or reserves exactly one coordinator run.
- That coordinator run executes the same planning and workflow path as a manually-started coordinator run. Resolved approval/autopilot settings determine whether confirmation can be unattended; pickup alone is not approval.
Backlog Task Lifecycle
The persisted task states are Backlog -> Ready -> Claimed. Running, completed, and failed are board projections of the linked run, not additional BacklogTaskState values. Use the shared backlog board explanation rather than a second lifecycle diagram.
The critical operation is the transition from Ready to Claimed. It must be atomic. If two heartbeat ticks or processes see the same ready task, only one should reserve the task and create the coordinator run. Otherwise, the system would execute duplicate plans for the same backlog item.
Heartbeat Loop
The heartbeat loop is intentionally simple and repeatable:
- Scan active projects.
- Skip projects whose workspace is unavailable.
- Read a deterministic top-N set of Ready tasks per project.
- For each task, attempt an atomic claim and run reservation.
- Start the reserved coordinator run under its resolved approval/autopilot settings.
- After the project loop, run one coordinator reconciliation sweep and drain orphaned OutcomeSpec decisions.
- Every configured Nth tick, run the optional AgentHost orphan-pod reaper.
The reconciliation, deferred-decision drain, and reaper are separate guarded phases outside the per-project loop (apps/Agentweaver.Api/Coordinator/CoordinatorHeartbeatService.cs:151–210).
Workflow overrides are allowed at the backlog task level, subject to registry availability and binding—not invocation-kind or trigger-eligibility filtering. Event and schedule producers initiate backlog work upstream; selection considers the valid available workflow set.
Why Heartbeat Instead of Immediate Execution?
A heartbeat loop gives the system backpressure and recovery:
- Projects can limit how many ready tasks are picked up per tick.
- Workspace availability can be checked before work starts.
- If the process crashes, unclaimed Ready tasks remain visible for the next tick.
- Claimed tasks can be reconciled against their reserved runs.
- The same mechanism can eventually support multiple workers if claim semantics stay atomic.
Where this lives:
apps/Agentweaver.Api/Coordinator/packages/Agentweaver.Domain/
Workflow Gates and Merge
Workflow-declared review gates
Review gates are declared in workflow nodes and edges. RunWorkflowFactory.ResolveEffectiveWorkflowAsync resolves a workflow and returns it without composing a separate project policy (apps/Agentweaver.Api/Runs/RunWorkflowFactory.cs:1495–1517). Blueprint validation accepts only review_policy: default (apps/Agentweaver.Api/Blueprints/BlueprintService.cs:113–115). Legacy policy-prefixed adapters are binding plumbing, not a registry or composer.
See binding declarative nodes to runtime execution for fail-closed gate/edge binding. Collective assembly executes the authored aggregate gates; a collective RAI RED verdict opens a durable human-review escalation (InReview / AwaitingReview, reason rai_red), not a terminal RaiBlocked dead end (apps/Agentweaver.Api/Coordinator/CoordinatorAssemblyService.cs:3752–3795).
Human Review as a Pause Point
Human review is not just an event; it is a pause in the workflow. The runtime emits a review request, the watch loop persists the run as awaiting review, and the stream can close cleanly while the system waits for user action.
The user action then chooses a path:
- approve and continue to merge,
- request changes and loop back to agent work,
- or decline and terminate.
This design keeps review durable and externally controllable. A browser tab can close while a run waits for review; the run state still tells the next client exactly what is needed.
The client submits a decision to the API; it does not resume a workflow directly. The shared review sequence owns the authorization, pending-request arbitration, and replay behavior.
Merge Gate
Merge is a gate because generated work can be correct but not mergeable. The merge step surfaces conflicts, blocked policies, or repository constraints.
A healthy merge gate should distinguish:
- blocked but recoverable — return to review or revision with a clear reason,
- merged — terminal success and scribe recording,
- terminal merge failure — cannot proceed without manual intervention.
The parent coordinator run owns merge for coordinated work. Child runs should not merge because they do not know whether sibling subtasks are complete or consistent.
Scribe
Scribe is the post-outcome memory step. It records what happened, decisions, learnings, or trace information after the run reaches the appropriate terminal path. Conceptually, Scribe turns execution history into reusable project memory.
Where this lives:
apps/Agentweaver.Api/Workflows/apps/Agentweaver.Api/Runs/
Recovery and Failure Handling
Agentweaver recovery is built from several smaller guarantees rather than one global transaction.
Idempotent Planning
If a coordinator run already has a WorkPlan, the coordinator should not decompose again. This prevents duplicate children and preserves the original confirmed intent.
Atomic Pickup
Backlog pickup should claim the task and reserve the run in one atomic operation. If reservation fails, the task should not appear successfully claimed without an executable run.
Durable Events
Events should be appended durably before live publication. This lets clients reconnect and lets operators inspect what happened after a crash.
Watch-Loop Status Projection
The runtime graph emits events. The watch loop projects those events into durable statuses. Keeping this projection centralized prevents every executor from inventing its own status semantics.
Frontier-Based Dispatch
The dispatcher can be rerun safely because it reads persisted subtask states and dependencies. Already-dispatched or completed subtasks are skipped; newly-ready pending subtasks can be launched.
Review and Merge Re-entry
Review and merge failures often are not terminal. A requested change loops back to agent work. A blocked merge can return to review. Only explicit terminal paths should mark the run failed, declined, merged, or merge-failed.
Reconciliation
A reconciler should periodically compare plans, subtasks, child runs, and parent status. Its job is to notice mismatches such as:
- a subtask marked running whose child run reached a terminal state,
- a plan whose all subtasks are assemble-ready but parent assembly has not started,
- a claimed backlog task whose reserved run was not launched,
- or a coordinator parent waiting on children that no longer exist.
The reconciler is what turns persisted state into eventual progress after crashes or partial failures.
Casting and Blueprints Integration
Casting provides the roster: agent names, role charters, default models, and required system agents such as Coordinator, Scribe, Ralph, and Rai. Orchestration consumes this roster when assigning subtasks and binding review responsibilities.
Blueprints provide defaults: initial roster, workflow set, default workflow, sandbox profile, and optional bespoke roles. The required review_policy field accepts only default; it does not configure additional injected gates. Applying a blueprint can materialize workflow definitions and persist defaults that later coordinator runs select from.
The key boundary is that Casting and Blueprints define who is available and what defaults apply. The orchestration engine decides what work is needed now and how that work moves through gates.
See team-casting.md for the detailed model.
Where this lives:
apps/Agentweaver.Api/Casting/apps/Agentweaver.Api/Blueprints/
Extension Points and Gotchas
- Do not treat workflow ids as executable code. A workflow must be parsed, classified, bound to known executors, and validated before it can run.
- Trigger evaluation is an ingress boundary. Verified events and schedules may initiate backlog work; they do not filter selector candidates by run origin.
- Child pipelines are intentionally shorter. Per-child review, merge, and scribe would fragment responsibility. Keep those phases at the parent level for coordinated work.
- Advisory isolation is not a lock. File ownership hints help dispatch and planning, but dependency edges, review, and merge conflict handling still matter.
- Workflow gate binding must fail closed. An unsupported declared gate or transition prevents execution rather than silently weakening the authored graph.
- Registry sync matters. Explicit sync gives immediate validation feedback; signature changes also refresh cached workflow results on the next read.
- Live streams and durable streams serve different users. Live channels make the UI responsive; durable event logs make reconnect and crash recovery possible. Keep both.
- Comments can drift from behavior. Prefer the persisted contracts and current service flow over historical comments when validating orchestration behavior.
Rebuilding Checklist
If you were rebuilding Agentweaver orchestration from scratch, implement in this order:
- Durable run records, statuses, and event log.
- Workflow definitions with separate trigger evaluation and fail-closed binding.
- Agent execution wrapped by a watch loop that projects events into statuses.
- Workflow-declared review gates with durable decisions.
- OutcomeSpec confirmation flow.
- WorkPlan, subtask, and dependency persistence.
- Frontier-based child dispatch and assemble-ready handoff.
- Parent assembly, review, merge, and scribe phases.
- Backlog Ready-to-Claimed atomic pickup.
- Heartbeat scanning and reconciliation.
- Casting and Blueprint defaults feeding coordinator selection.
The central design principle is simple: persist intent, execute only eligible work, make every gate explicit, and recover by replaying durable state rather than reinterpreting the original request.
Diagram details and constraints
| Element | Contract |
|---|---|
| title | Generic default workflow |
| subtitle | Built-in template • merge → PR publication → Scribe |
| returns-heading | SOURCE / RETURN |
| outcomes-heading | OUTCOMES |
| footer | PR action can skip / fail and still reach Scribe. No-changes also reaches Scribe. |
| Agent work | Agent |
| Agent work | Agent task |
| Agent work | agent |
| RAI gate | Rai |
| RAI gate | Verdict routing |
| RAI gate | rai |
| Human review | Review |
| Human review | human-review |
| Merge | Merge |
| Merge | Merge outcome routing |
| Merge | merge |
| Publish / reuse PR | Publish / reuse PR |
| Publish / reuse PR | Create / reuse; not git push |
| Publish / reuse PR | action |
| Scribe | Scribe |
| Scribe | Record the run outcome |
| Scribe | scribe |
| Safety failed | Safety failed |
| Safety failed | Workflow endpoint |
| Declined | Declined |
| Done | Done |
| edge-02-label | revise |
| edge-03-label | safety- failed |
| edge-04-label | no- changes |
| edge-05-label | review |
| edge-06-label | approved |
| edge-07-label | request-changes |
| edge-08-label | declined |
| edge-09-label | merged |
| edge-10-label | blocked |
Diagram details and constraints
| Element | Contract |
|---|---|
| title | Postgres is the event relay |
| subtitle | Any API replica can serve a cursor over durable RunEvents—no sticky session required. |
| group-title0 | Write path · replica A |
| group-title1 | Read path · replica B |
| Run producer | Run producer |
| Run producer | Append a structured event |
| Run producer | runId + type + payload |
| EF event stream | EF event stream |
| EF event stream | Serialize writes per run |
| EF event stream | pg_advisory_xact_lock |
| RunEvents | RunEvents |
| RunEvents | Shared PostgreSQL table |
| RunEvents | (RunId, Sequence) |
| Web / MCP watcher | Web / MCP watcher |
| Web / MCP watcher | Consume ordered events |
| Web / MCP watcher | last delivered cursor |
| SSE endpoint | SSE endpoint |
| SSE endpoint | Emit id + event + data |
| SSE endpoint | ordered response frames |
| EF subscriber | EF subscriber |
| EF subscriber | Read Sequence > cursor |
| EF subscriber | idle poll: 250 ms |
| e1 | append |
| e2 | commit |
| e3 | ordered batch |
| e4 | yield |
| e5 | SSE frames |
| assurance-title | POSTGRES LANE ONLY |
| assurance-line1 | SQLite register-channel / replay / tail is a separate implementation—not this architecture. |
| assurance-line2 | Late-delta suppression is process-local; do not read it as a database-wide terminal fence. |
| Run producer | Input |
| Run producer | RunStreamEntry |
| Run producer | Identity |
| Run producer | runId + event type |
| Run producer | Body |
| Run producer | Structured payload |
| Run producer | Ack |
| Run producer | After durable commit |
| EF event stream | Lock |
| EF event stream | Per-run advisory lock |
| EF event stream | Next |
| EF event stream | MAX(Sequence) + 1 |
| EF event stream | Write |
| EF event stream | Save transaction |
| EF event stream | Commit |
| EF event stream | Before acknowledgement |
| RunEvents | Table |
| RunEvents | Key |
| RunEvents | RunId + Sequence |
| RunEvents | Order |
| RunEvents | Ascending sequence |
| RunEvents | Reuse |
| RunEvents | Same type / payload |
| Web / MCP watcher | Client |
| Web / MCP watcher | Web or MCP |
| Web / MCP watcher | Resume |
| Web / MCP watcher | Last delivered cursor |
| Web / MCP watcher | Replica |
| Web / MCP watcher | No sticky requirement |
| Web / MCP watcher | History |
| Web / MCP watcher | Durable ordered events |
| SSE endpoint | Frame |
| SSE endpoint | id + event + data |
| SSE endpoint | Cursor |
| SSE endpoint | Last-Event-ID |
| SSE endpoint | Delivery |
| SSE endpoint | Yield ordered events |
| SSE endpoint | Close |
| SSE endpoint | After batch is drained |
| EF subscriber | Query |
| EF subscriber | Sequence > cursor |
| EF subscriber | Idle |
| EF subscriber | Poll after 250 ms |
| EF subscriber | State |
| EF subscriber | Shared durable table |
| EF subscriber | Blocked |
| EF subscriber | Retryable: keep open |
| producer | Coordinator or run execution; Acknowledgement follows commit |
| append | Allocate MAX(Sequence) + 1; Save and commit transaction |
| store | Cross-replica ordered history; Explicit duplicates must match payload |
| client | Reconnect from the cursor; No local channel dependency |
| sse | Cursor advances after delivery; Drain batch before terminal close |
| reader | Query the shared durable table; Retryable assembly_blocked stays open |
| notes | POSTGRES LANE ONLY; SQLite register-channel / replay / tail is a separate implementation—not this architecture.; Late-delta suppression is process-local; do not read it as a database-wide terminal fence. |
| groups | Write path · replica A; Read path · replica B |
Diagram details and constraints
| Element | Contract |
|---|---|
| title | One heartbeat tick, two scopes |
| takeaway | Pickup runs per project; reconciliation, deferred-spec drain and optional reaping run afterward. |
| group-title-0 | PROJECT LOOP |
| group-title-1 | PER-PROJECT PICKUP |
| group-title-2 | ONCE AFTER THE PROJECT LOOP |
| Heartbeat tick | Heartbeat tick |
| Heartbeat tick | Enumerate projects |
| Heartbeat tick | failure-isolated sweep |
| Active + available? | Active + available? |
| Active + available? | Skip unavailable projects |
| Active + available? | per-project admission |
| Capped Ready list | Capped Ready list |
| Capped Ready list | Deterministic candidates |
| Capped Ready list | per-project limit |
| Atomic claim | Atomic claim |
| Atomic claim | Reserve coordinator run |
| Atomic claim | competing claim may lose |
| Start reserved run | Start reserved run |
| Start reserved run | Carry backlog origin |
| Start reserved run | confirmation policy applies |
| End project loop | End project loop |
| End project loop | Record tick result |
| End project loop | not an inner-loop sweep |
| Reconcile once | Reconcile once |
| Reconcile once | Repair durable supervision |
| Reconcile once | after all projects |
| Drain spec decisions | Drain spec decisions |
| Drain spec decisions | Recover orphaned decisions |
| Drain spec decisions | durable OutcomeSpec |
| Optional pod reaper | Optional pod reaper |
| Optional pod reaper | Every N ticks when enabled |
| Optional pod reaper | throttled cleanup |
| e0 | each |
| e1 | eligible |
| e2 | claim |
| e3 | won |
| e4 | loop done |
| e5 | once |
| e6 | then |
| e7 | when due |
| groups | PROJECT LOOP; PER-PROJECT PICKUP; ONCE AFTER THE PROJECT LOOP |
Diagram details and constraints
| Element | Contract |
|---|---|
| title | A dependency DAG, not agent chat |
| takeaway | Illustrative tasks A-D show readiness; only assemble-ready/completed prerequisites count. |
| group-title-0 | DURABLE INTENT AND READINESS |
| group-title-1 | ILLUSTRATIVE PARALLEL ROOTS |
| group-title-2 | ILLUSTRATIVE DEPENDENTS AND HANDOFF |
| Confirmed OutcomeSpec | Confirmed OutcomeSpec |
| Confirmed OutcomeSpec | Intent before decomposition |
| Confirmed OutcomeSpec | confirmed status |
| Persisted WorkPlan | Persisted WorkPlan |
| Persisted WorkPlan | Tasks + dependency edges |
| Persisted WorkPlan | selected workflow |
| Readiness rule | Readiness rule |
| Readiness rule | Every predecessor satisfied |
| Readiness rule | assemble_ready / done |
| Example root A | Example root A |
| Example root A | No prerequisites |
| Example root A | illustrative, not fixed |
| Example root B | Example root B |
| Example root B | parallel with A |
| Satisfied roots | Satisfied roots |
| Satisfied roots | Not merely terminal |
| Satisfied roots | failure does not unlock |
| Example dependent C | Example dependent C |
| Example dependent C | Depends on A |
| Example dependent C | illustrative edge |
| Example dependent D | Example dependent D |
| Example dependent D | Depends on A and B |
| Example dependent D | illustrative join |
| Collective handoff | Collective handoff |
| Collective handoff | Recheck aggregate eligibility |
| Collective handoff | quiescence != success |
| e0 | persist |
| e1 | evaluate |
| e2 | ready |
| e4 | A done |
| e6 | B done |
| e7 | both |
| e8 | settled |
| groups | DURABLE INTENT AND READINESS; ILLUSTRATIVE PARALLEL ROOTS; ILLUSTRATIVE DEPENDENTS AND HANDOFF |
Diagram details and constraints
| Element | Contract |
|---|---|
| title | Child execution and dependency progress |
| takeaway | Children publish typed results and branch content; collective review belongs to the parent. |
| group-title-0 | DEPENDENCY ADMISSION |
| group-title-1 | TRIMMED CHILD EXECUTION |
| group-title-2 | PUBLISHED CONTENT AND AGGREGATE ELIGIBILITY |
| Pending subtasks | Pending subtasks |
| Pending subtasks | Persisted dependency map |
| Pending subtasks | not automatically ready |
| Ready frontier | Ready frontier |
| Ready frontier | All predecessors satisfied |
| Ready frontier | not any terminal result |
| Isolated child | Isolated child |
| Isolated child | Launch agent execution |
| Isolated child | separate checkout |
| Agent result | Agent result |
| Agent result | Typed conditional output |
| Agent result | no child RAI executor |
| Assemble-ready | Assemble-ready |
| Assemble-ready | Successful child content |
| Assemble-ready | satisfies dependents |
| Typed turn failure | Typed turn failure |
| Typed turn failure | Does not satisfy dependents |
| Typed turn failure | no per-child review |
| Published branch | Published branch |
| Published branch | Authoritative committed tip |
| Published branch | not shared mutable files |
| Dependency base | Dependency base |
| Dependency base | Rebuild prerequisite content |
| Dependency base | new isolated dependent |
| Parent assembly check | Parent assembly check |
| Parent assembly check | Quiescence plus eligibility |
| Parent assembly check | collective gates later |
| e0 | evaluate |
| e1 | dispatch |
| e2 | execute |
| e3 | success |
| e4 | failed |
| e5 | publish |
| e6 | integrate |
| e7 | unlock |
| e8 | settled |
| e9 | blocked |
| groups | DEPENDENCY ADMISSION; TRIMMED CHILD EXECUTION; PUBLISHED CONTENT AND AGGREGATE ELIGIBILITY |
Diagram details and constraints
| Element | Contract |
|---|---|
| title | Persist the plan before dispatch |
| takeaway | Confirmation, selection and decomposition precede durable child dispatch. |
| group-title-0 | INTENT AND REUSE |
| group-title-1 | SELECTION AND DURABLE PLAN |
| group-title-2 | DISPATCH AND COLLECTIVE HANDOFF |
| Submitted request | Submitted request |
| Submitted request | Goal + caller context |
| Submitted request | manual or pickup |
| Confirmation boundary | Confirmation boundary |
| Confirmation boundary | Manual or unattended policy |
| Confirmation boundary | autopilot-dependent |
| Existing plan? | Existing plan? |
| Existing plan? | Reuse persisted plan |
| Existing plan? | avoid decomposing twice |
| Select workflow | Select workflow |
| Select workflow | Available definitions |
| Select workflow | explicit choices honored |
| Decompose + validate | Decompose + validate |
| Decompose + validate | Outcome-complete work |
| Decompose + validate | compatibility check |
| Persist WorkPlan | Persist WorkPlan |
| Persist WorkPlan | Subtasks and dependencies |
| Persist WorkPlan | workflow identity |
| Ready frontier | Ready frontier |
| Ready frontier | Dependency satisfaction |
| Ready frontier | pending -> ready work |
| Dispatch children | Dispatch children |
| Dispatch children | Observe classified outcomes |
| Dispatch children | isolated child runs |
| Collective handoff | Collective handoff |
| Collective handoff | After child supervision |
| Collective handoff | assembly eligibility |
| e0 | confirm |
| e1 | lookup |
| e2 | new |
| e3 | decompose |
| e4 | persist |
| e5 | reuse |
| e6 | ready |
| e7 | dispatch |
| e8 | handoff |
| groups | INTENT AND REUSE; SELECTION AND DURABLE PLAN; DISPATCH AND COLLECTIVE HANDOFF |
Diagram details and constraints
| Element | Contract |
|---|---|
| title | Execution and observation cooperate |
| takeaway | The watcher projects runtime events into durable state; it is not the executing graph. |
| group-title-0 | START AND BIND |
| group-title-1 | RUNTIME AND SUPERVISION |
| group-title-2 | DURABLE PROJECTIONS |
| Run orchestrator | Run orchestrator |
| Run orchestrator | Starts factory + watcher |
| Run orchestrator | supervised lifetime |
| Effective definition | Effective definition |
| Effective definition | Resolve concrete workflow |
| Effective definition | no policy composer |
| Factory + binder | Factory + binder |
| Factory + binder | Build executable graph |
| Factory + binder | typed bindings |
| Checkpointed stream | Checkpointed stream |
| Checkpointed stream | MAF executes the graph |
| Checkpointed stream | provider-aware store |
| Watch loop | Watch loop |
| Watch loop | Consumes runtime updates |
| Watch loop | not graph execution |
| Review request | Review request |
| Review request | Persist pending decision |
| Review request | durable pause context |
| Typed terminal | Typed terminal |
| Typed terminal | Classify completed output |
| Typed terminal | not inferred from text |
| Durable run state | Durable run state |
| Durable run state | Persist status projection |
| Durable run state | watcher owns updates |
| Workflow-step events | Workflow-step events |
| Workflow-step events | Expose execution progress |
| Workflow-step events | client observation |
| e0 | start |
| e1 | resolve |
| e2 | execute |
| e3 | supervise |
| e4 | stream |
| e5 | request |
| e6 | terminal |
| e7 | persist |
| e8 | publish |
| groups | START AND BIND; RUNTIME AND SUPERVISION; DURABLE PROJECTIONS |
Diagram details and constraints
| Element | Contract |
|---|---|
| title | Review API decision paths |
| takeaway | Authorize first. Deliver through the right path. Lock before any merge CAS. |
| group-title-0 | ADMISSION AND REPLAY |
| group-title-1 | DELIVERY ALTERNATIVES |
| group-title-2 | CONTINUATION AND MERGE |
| Caller + access | Caller + access |
| Caller + access | Project contributor check |
| Caller + access | legacy: pending owner |
| Reviewable state? | Reviewable state? |
| Reviewable state? | Inspect status + pending |
| Reviewable state? | awaiting_review |
| Replay or conflict | Replay or conflict |
| Replay or conflict | Matching terminal: reuse |
| Replay or conflict | otherwise: 409 |
| Live pending | Live pending |
| Live pending | Changes / decline use CAS |
| Live pending | approve: no merge CAS |
| Deferred pending | Deferred pending |
| Deferred pending | Persist the decision first |
| Deferred pending | then status transition |
| No live / no pending | No live / no pending |
| No live / no pending | Validate direct approval |
| No live / no pending | changes: 409 |
| Consume + deliver | Consume + deliver |
| Consume + deliver | Send workflow response |
| Consume + deliver | live continuation |
| Repository lock | Repository lock |
| Repository lock | Only on reaching merge |
| Repository lock | lock before CAS |
| Merge CAS + Git | Merge CAS + Git |
| Merge CAS + Git | Guard reviewed tree input |
| Merge CAS + Git | release lock on exit |
| e0 | check |
| e1 | replay |
| e2 | live |
| e3 | deferred |
| e4 | direct |
| e5 | deliver |
| e6 | on merge |
| e7 | approve |
| e8 | locked |
| groups | ADMISSION AND REPLAY; DELIVERY ALTERNATIVES; CONTINUATION AND MERGE |
