Skip to content

Orchestration Engine — Conceptual Deep Dive ​

Purpose & Mental Model ​

Agentweaver orchestration answers one question: how does a high-level goal become safe, reviewable, mergeable work performed by a team of agents?

The engine is intentionally split into two layers:

  1. Coordinator orchestration decides what should happen. It turns an ambiguous goal or backlog item into a confirmed outcome, decomposes that outcome into a dependency-aware plan, assigns work to team members, and assembles the results.
  2. Run workflow orchestration decides how each run moves through gates. It applies a declarative workflow to live execution: agent work, safety review, human review, merge, and scribe recording.

The split matters. Planning and decomposition need durable state, idempotency, and team-level reasoning. Individual run execution needs streaming, review gates, restart loops, and terminal status handling. Keeping those concerns separate lets Agentweaver recover from partial progress without re-asking the model to re-invent the plan.

A useful rebuilding rule is: the coordinator owns intent and coordination; workflows own execution gates.

Casting and Blueprints feed orchestration with team shape, role charters, and workflow defaults. Blueprint validation accepts only review_policy: default; this is not a configurable project review-policy overlay. They are summarized here only; the detailed explanation lives in team-casting.md.

Live workflow execution and non-workflow single-prompt paths use the Copilot SDK. The live workflow worker path does not switch worker implementation based on the run's model source; custom providers are configured through the SDK.

Core Design Invariants ​

These invariants are the backbone of the system:

  • Persist decisions before doing work. The requested outcome and work plan are stored before child runs are launched. Recovery starts from persisted intent, not from chat history.
  • Confirm ambiguity at the boundary. The coordinator may draft, revise, and ask for confirmation before committing a plan. Once confirmed, later components can assume the outcome is intentional.
  • Use declarative graphs for policy. Workflows describe nodes, gates, and edges. Runtime code binds those declarations to executable steps and fails closed when a step cannot be safely bound.
  • Advance only the ready frontier. Subtasks form a DAG. A subtask can run only after its dependencies are complete, so parallelism is safe and deterministic.
  • Separate child work from collective responsibility. Child runs produce pieces ready for assembly, not independently reviewed pieces. The parent coordinator assembles, reviews, merges, and records the combined outcome.
  • Make gates explicit and durable. Safety, human review, merge, and terminal states are visible run states and stream events, not hidden control flow.
  • Prefer idempotent recovery over clever replay. If a plan already exists, reuse it. If a run already reached a gate, resume from that gate. If a stream disconnects, replay durable events.

Coordinator Orchestration ​

Problem It Solves ​

A user often gives Agentweaver a goal, not a task list. The coordinator converts that goal into something a team can execute safely:

  • What exact outcome are we trying to produce?
  • What assumptions or constraints define success?
  • Which parts can run independently?
  • Which specialist should own each part?
  • What must be reviewed before changes merge?

Without this layer, every agent run would independently interpret the same broad request. That leads to duplicated work, conflicting edits, and unclear ownership.

OutcomeSpec: The Intent Contract ​

The first durable artifact is the OutcomeSpec. Conceptually, it is the contract between the requester and the system.

It captures:

  • the original goal,
  • the desired outcome,
  • scope and exclusions,
  • assumptions,
  • clarifying questions or revision feedback,
  • and whether the outcome is still being drafted, awaiting confirmation, confirmed, or declined.

The important design choice is that confirmation happens before decomposition is treated as authoritative. The coordinator can draft an interpretation, receive revision feedback, and loop until the requester or unattended policy confirms it.

Because the confirmed OutcomeSpec is the source of truth for intent, it — not the later decomposition — is where the work's breadth originates. When the drafter frames the outcome it reads the project's team roster and honors the breadth the goal explicitly asks for: a full-lifecycle "from the initial idea through to a working app" goal yields an outcome whose scope enumerates the discovery, PM, design, and build deliverables the goal warrants (filtered to what the team can actually produce), while a narrow goal stays lean. Downstream workflow selection and decomposition then faithfully honor that confirmed breadth, so if the outcome is narrowed at drafting time, the whole PM/discovery half is silently dropped everywhere after it. See How drafting works for the roster-awareness and scope-breadth guidance that shapes the draft.

Rebuild guidance: treat the OutcomeSpec as the source of truth for intent. Do not let individual worker agents reinterpret the original request independently once the spec is confirmed.

WorkPlan: The Execution Contract ​

After confirmation, the coordinator creates a WorkPlan. The WorkPlan is the execution contract for the parent coordinator run.

It stores:

  • the confirmed OutcomeSpec it implements,
  • the selected workflow,
  • subtask records,
  • dependency edges between subtasks,
  • assembly state,
  • and any integration branch or coordination metadata.

Each subtask includes its assigned agent, model choice, charter/context, isolation intent, status, child run id, and any recovery guidance.

The plan is a DAG because ordering is a correctness constraint. If subtask B depends on subtask A, B should not start merely because an agent is free. This allows safe parallelism: every tick can dispatch all currently-ready nodes while preserving required sequencing.

Rebuild guidance: store the plan before dispatch. If the coordinator crashes after planning but before child runs start, it should resume from the persisted WorkPlan rather than ask a model to decompose again.

Coordinator Control Flow ​

The coordinator flow has two phases:

  1. Model-assisted planning phase — draft and confirm the OutcomeSpec, select a workflow, decompose the work, and persist the WorkPlan.
  2. Service-driven execution phase — dispatch ready subtasks, watch child runs, assemble results, and advance the parent run through review and merge gates.

The coordinator is designed to be idempotent. If it is asked to orchestrate a run that already has a WorkPlan, it does not create a second plan. That invariant prevents duplicate child runs and conflicting DAGs.

Decomposition Logic ​

A good decomposition algorithm should produce subtasks that are:

  • owned by one agent,
  • bounded enough to complete independently,
  • ordered by explicit dependencies,
  • labeled with intended isolation or file ownership,
  • and recoverable with enough guidance to retry or inspect failures.

Agentweaver treats dependency edges and file/isolation hints as coordination data. The dependency graph is the hard ordering rule. Isolation hints are advisory: they help avoid conflicts and guide dispatch, but they are not a substitute for merge conflict handling or review.

Cycle breaking is essential. Model-generated plans can accidentally create circular dependencies. A production coordinator should detect cycles and either remove weak edges, ask for clarification, or fail before dispatch. Dispatching a cyclic plan would deadlock because no frontier can become ready.

Dispatch and Assembly ​

The dispatcher repeatedly asks: which pending subtasks have all dependencies completed? Those subtasks form the ready frontier.

For each ready subtask, it launches a child run with an isolated working tree and output branch. Child runs are intentionally trimmed: they perform agent work, then stop at an assemble-ready boundary. They do not each perform RAI, human review, merge, or scribe. Those are parent-level responsibilities because the user reviews the combined outcome, not a pile of isolated fragments.

When a child reaches assemble-ready/completed, the dispatcher rebuilds the coordinator integration branch from the successful child branches in dependency order. Dependents are then branched from that integration branch, so they can read files produced by their prerequisites without concurrent siblings sharing one mutable git index.

Assembly is where the coordinator turns independent child outputs into one coherent result. This is also where conflicts, missing pieces, and cross-subtask inconsistencies should be detected before the parent enters review and merge gates.

The child graph branches on AgentTurnOutput.TerminalFailureReason: a clean turn reaches child-assemble-ready, while a typed failure reaches child-turn-failed (apps/Agentweaver.Api/Runs/RunWorkflowFactory.cs:788–814). “All subtasks settled” is not equivalent to “all outputs eligible for assembly”; a failed child is not an approved aggregate input.

Where this lives:

  • apps/Agentweaver.Api/Coordinator/
  • apps/Agentweaver.Api/Memory/

Workflows and Trigger Evaluation ​

Workflow as Policy Graph ​

A workflow is not just a list of functions. It is a policy graph that describes how a run should progress through work, checks, review, merge, and terminal states.

A workflow definition answers:

  • What starts the graph?
  • Which node performs agent work?
  • Which gates can send work back for revision?
  • Which failures are terminal?
  • Which path means success?
  • Which event or schedule declarations can initiate backlog work for this workflow?

The shared illustration describes the standalone built-in workflow, not mandatory collective assembly policy. Its current success path is agent -> rai -> review -> merge -> push-pr -> scribe -> done (apps/Agentweaver.Api/Workflows/DefaultWorkflowTemplate.cs:42–161); child runs bypass this graph.

The important idea is that loops are first-class. Safety or review can return work to the producer. Merge can return to review if blocked. Terminal failures are explicit exits, not exceptions swallowed by the runtime.

Invocation context and event triggers ​

Manual and heartbeat origins are recorded as invocation context. They do not remove valid workflows from the selector candidate set. Event and schedule triggers are evaluated before the normal backlog and coordinator pickup path. A requested override is used only when it resolves to a valid, bindable workflow.

Workflow Selection Logic ​

The selection order is deliberately conservative:

  1. Load built-in, catalog/library, and project-authored workflows.
  2. Record invocation kind from the run origin.
  3. Honor a valid override if present.
  4. Order the configured project default first without short-circuiting automatic selection.
  5. If exactly one workflow remains, use it without model help.
  6. If several remain, ask the selector to choose the best process fit.
  7. If selector output is invalid or parsing fails, fall back safely rather than inventing a workflow id.

This pattern limits model authority. The model may choose among safe candidates, but it cannot bypass validation or runtime binding.

See workflow selection for override events, selector retry/fallback behavior, and the post-decomposition Build & Test compatibility check. That check can choose a platform software workflow when automatic project candidates cannot cover code-producing work; an explicit workflow lacking Build & Test is honored with a warning.

Binding Declarative Nodes to Runtime Execution ​

A workflow file describes intent. The runtime must bind that intent to concrete executors.

The binder should:

  • classify nodes by type and gate kind,
  • resolve each node to a known executor,
  • expand logical edges into the live execution graph,
  • verify every workflow-declared gate and transition has a binding,
  • and fail closed if a required node cannot be executed safely.

Failing closed is a security and correctness property. A workflow that asks for a safety gate but cannot bind one should not silently skip safety. Likewise, a custom node type should not become a no-op merely because the binder does not understand it.

Some workflow shapes, such as fan-out/fan-in style nodes, are design-level extension points: the graph model can express them, and the binder is the place where their executors are resolved. They give the system room to grow more complex execution patterns without changing the surrounding contract.

Where this lives:

  • apps/Agentweaver.Api/Workflows/
  • docs/workflow-library.md
  • docs/workflow-binder.md

Run Lifecycle ​

What a Run Represents ​

A run is the durable unit of execution. Conceptually it bundles:

  • a project and workspace/worktree,
  • the assigned agent and charter/context,
  • the selected workflow,
  • the run origin,
  • live and durable event streams,
  • and a persisted status.

A run can be started directly by a user, reserved by backlog pickup, created as a coordinator parent, or launched as a coordinator child. All forms should converge on the same lifecycle machinery so status, streaming, review, and recovery behave consistently.

Parent, Child, and Pickup Runs ​

Agentweaver uses run origin to preserve intent:

  • Manual runs are user-started and usually go through the full workflow.
  • Coordinator parent runs own the team-level plan, assembly, review, merge, and scribe phases.
  • Coordinator child runs execute one subtask and stop at the assemble-ready boundary after agent work.
  • Backlog pickup runs are coordinator runs created by the heartbeat loop for unattended ready tasks.

The key difference is not the storage shape; it is the responsibility boundary. Child runs should not merge independently because they are fragments of the parent outcome. Parent runs should not redo child work because they coordinate, assemble, and gate the whole result.

State Machine ​

This is a conceptual, non-exhaustive state machine. It emphasizes externally visible gates. When a human review node is reached, the run becomes AwaitingReview and the client can act. When merge is requested, the run becomes Merging. These are not merely internal events; they are durable states used by clients, recovery, and monitoring.

Runtime Sequence ​

The watch loop translates live runtime events into persisted run state. This keeps state transitions centralized. The agent produces work; the workflow emits events; the watch loop decides what those events mean for durable status and client-visible stream completion.

Event Streaming ​

Run events have two purposes:

  1. Live feedback — clients can see what the agent is doing now.
  2. Recovery and reconnect — clients can replay what happened if they disconnect or the process restarts.

The Postgres path appends durably, then reads ordered rows after the subscriber's cursor:

EfRunEventStream allocates the next sequence under a per-run advisory transaction lock and acknowledges only after commit. Subscribers query Sequence > cursor and poll again after 250 ms when no rows are available. A reconnect can therefore land on another API replica without relying on the first replica's channel. The SQLite/local alternative has a bounded process-local channel; that channel is not the Postgres cross-replica delivery mechanism. The SSE endpoint supplies framing and completion behavior.

Where this lives:

  • apps/Agentweaver.Api/Runs/
  • packages/Agentweaver.AgentRuntime/Workflow/
  • apps/Agentweaver.Api/Infrastructure/
  • docs/run-event-stream.md

Backlog and Heartbeat Pickup ​

Problem It Solves ​

The backlog lets Agentweaver accept work before an agent is actively assigned. The heartbeat loop turns ready backlog items into unattended coordinator runs.

This separates commitment from execution:

  • A task can be captured and ordered in the backlog.
  • Later, when it becomes ready and workspace conditions allow, the system claims it.
  • Claiming creates or reserves exactly one coordinator run.
  • That coordinator run executes the same planning and workflow path as a manually-started coordinator run. Resolved approval/autopilot settings determine whether confirmation can be unattended; pickup alone is not approval.

Backlog Task Lifecycle ​

The persisted task states are Backlog -> Ready -> Claimed. Running, completed, and failed are board projections of the linked run, not additional BacklogTaskState values. Use the shared backlog board explanation rather than a second lifecycle diagram.

The critical operation is the transition from Ready to Claimed. It must be atomic. If two heartbeat ticks or processes see the same ready task, only one should reserve the task and create the coordinator run. Otherwise, the system would execute duplicate plans for the same backlog item.

Heartbeat Loop ​

The heartbeat loop is intentionally simple and repeatable:

  1. Scan active projects.
  2. Skip projects whose workspace is unavailable.
  3. Read a deterministic top-N set of Ready tasks per project.
  4. For each task, attempt an atomic claim and run reservation.
  5. Start the reserved coordinator run under its resolved approval/autopilot settings.
  6. After the project loop, run one coordinator reconciliation sweep and drain orphaned OutcomeSpec decisions.
  7. Every configured Nth tick, run the optional AgentHost orphan-pod reaper.

The reconciliation, deferred-decision drain, and reaper are separate guarded phases outside the per-project loop (apps/Agentweaver.Api/Coordinator/CoordinatorHeartbeatService.cs:151–210).

Workflow overrides are allowed at the backlog task level, subject to registry availability and binding—not invocation-kind or trigger-eligibility filtering. Event and schedule producers initiate backlog work upstream; selection considers the valid available workflow set.

Why Heartbeat Instead of Immediate Execution? ​

A heartbeat loop gives the system backpressure and recovery:

  • Projects can limit how many ready tasks are picked up per tick.
  • Workspace availability can be checked before work starts.
  • If the process crashes, unclaimed Ready tasks remain visible for the next tick.
  • Claimed tasks can be reconciled against their reserved runs.
  • The same mechanism can eventually support multiple workers if claim semantics stay atomic.

Where this lives:

  • apps/Agentweaver.Api/Coordinator/
  • packages/Agentweaver.Domain/

Workflow Gates and Merge ​

Workflow-declared review gates ​

Review gates are declared in workflow nodes and edges. RunWorkflowFactory.ResolveEffectiveWorkflowAsync resolves a workflow and returns it without composing a separate project policy (apps/Agentweaver.Api/Runs/RunWorkflowFactory.cs:1495–1517). Blueprint validation accepts only review_policy: default (apps/Agentweaver.Api/Blueprints/BlueprintService.cs:113–115). Legacy policy-prefixed adapters are binding plumbing, not a registry or composer.

See binding declarative nodes to runtime execution for fail-closed gate/edge binding. Collective assembly executes the authored aggregate gates; a collective RAI RED verdict opens a durable human-review escalation (InReview / AwaitingReview, reason rai_red), not a terminal RaiBlocked dead end (apps/Agentweaver.Api/Coordinator/CoordinatorAssemblyService.cs:3752–3795).

Human Review as a Pause Point ​

Human review is not just an event; it is a pause in the workflow. The runtime emits a review request, the watch loop persists the run as awaiting review, and the stream can close cleanly while the system waits for user action.

The user action then chooses a path:

  • approve and continue to merge,
  • request changes and loop back to agent work,
  • or decline and terminate.

This design keeps review durable and externally controllable. A browser tab can close while a run waits for review; the run state still tells the next client exactly what is needed.

The client submits a decision to the API; it does not resume a workflow directly. The shared review sequence owns the authorization, pending-request arbitration, and replay behavior.

Merge Gate ​

Merge is a gate because generated work can be correct but not mergeable. The merge step surfaces conflicts, blocked policies, or repository constraints.

A healthy merge gate should distinguish:

  • blocked but recoverable — return to review or revision with a clear reason,
  • merged — terminal success and scribe recording,
  • terminal merge failure — cannot proceed without manual intervention.

The parent coordinator run owns merge for coordinated work. Child runs should not merge because they do not know whether sibling subtasks are complete or consistent.

Scribe ​

Scribe is the post-outcome memory step. It records what happened, decisions, learnings, or trace information after the run reaches the appropriate terminal path. Conceptually, Scribe turns execution history into reusable project memory.

Where this lives:

  • apps/Agentweaver.Api/Workflows/
  • apps/Agentweaver.Api/Runs/

Recovery and Failure Handling ​

Agentweaver recovery is built from several smaller guarantees rather than one global transaction.

Idempotent Planning ​

If a coordinator run already has a WorkPlan, the coordinator should not decompose again. This prevents duplicate children and preserves the original confirmed intent.

Atomic Pickup ​

Backlog pickup should claim the task and reserve the run in one atomic operation. If reservation fails, the task should not appear successfully claimed without an executable run.

Durable Events ​

Events should be appended durably before live publication. This lets clients reconnect and lets operators inspect what happened after a crash.

Watch-Loop Status Projection ​

The runtime graph emits events. The watch loop projects those events into durable statuses. Keeping this projection centralized prevents every executor from inventing its own status semantics.

Frontier-Based Dispatch ​

The dispatcher can be rerun safely because it reads persisted subtask states and dependencies. Already-dispatched or completed subtasks are skipped; newly-ready pending subtasks can be launched.

Review and Merge Re-entry ​

Review and merge failures often are not terminal. A requested change loops back to agent work. A blocked merge can return to review. Only explicit terminal paths should mark the run failed, declined, merged, or merge-failed.

Reconciliation ​

A reconciler should periodically compare plans, subtasks, child runs, and parent status. Its job is to notice mismatches such as:

  • a subtask marked running whose child run reached a terminal state,
  • a plan whose all subtasks are assemble-ready but parent assembly has not started,
  • a claimed backlog task whose reserved run was not launched,
  • or a coordinator parent waiting on children that no longer exist.

The reconciler is what turns persisted state into eventual progress after crashes or partial failures.

Casting and Blueprints Integration ​

Casting provides the roster: agent names, role charters, default models, and required system agents such as Coordinator, Scribe, Ralph, and Rai. Orchestration consumes this roster when assigning subtasks and binding review responsibilities.

Blueprints provide defaults: initial roster, workflow set, default workflow, sandbox profile, and optional bespoke roles. The required review_policy field accepts only default; it does not configure additional injected gates. Applying a blueprint can materialize workflow definitions and persist defaults that later coordinator runs select from.

The key boundary is that Casting and Blueprints define who is available and what defaults apply. The orchestration engine decides what work is needed now and how that work moves through gates.

See team-casting.md for the detailed model.

Where this lives:

  • apps/Agentweaver.Api/Casting/
  • apps/Agentweaver.Api/Blueprints/

Extension Points and Gotchas ​

  • Do not treat workflow ids as executable code. A workflow must be parsed, classified, bound to known executors, and validated before it can run.
  • Trigger evaluation is an ingress boundary. Verified events and schedules may initiate backlog work; they do not filter selector candidates by run origin.
  • Child pipelines are intentionally shorter. Per-child review, merge, and scribe would fragment responsibility. Keep those phases at the parent level for coordinated work.
  • Advisory isolation is not a lock. File ownership hints help dispatch and planning, but dependency edges, review, and merge conflict handling still matter.
  • Workflow gate binding must fail closed. An unsupported declared gate or transition prevents execution rather than silently weakening the authored graph.
  • Registry sync matters. Explicit sync gives immediate validation feedback; signature changes also refresh cached workflow results on the next read.
  • Live streams and durable streams serve different users. Live channels make the UI responsive; durable event logs make reconnect and crash recovery possible. Keep both.
  • Comments can drift from behavior. Prefer the persisted contracts and current service flow over historical comments when validating orchestration behavior.

Rebuilding Checklist ​

If you were rebuilding Agentweaver orchestration from scratch, implement in this order:

  1. Durable run records, statuses, and event log.
  2. Workflow definitions with separate trigger evaluation and fail-closed binding.
  3. Agent execution wrapped by a watch loop that projects events into statuses.
  4. Workflow-declared review gates with durable decisions.
  5. OutcomeSpec confirmation flow.
  6. WorkPlan, subtask, and dependency persistence.
  7. Frontier-based child dispatch and assemble-ready handoff.
  8. Parent assembly, review, merge, and scribe phases.
  9. Backlog Ready-to-Claimed atomic pickup.
  10. Heartbeat scanning and reconciliation.
  11. Casting and Blueprint defaults feeding coordinator selection.

The central design principle is simple: persist intent, execute only eligible work, make every gate explicit, and recover by replaying durable state rather than reinterpreting the original request.

Diagram details and constraints
ElementContract
titleGeneric default workflow
subtitleBuilt-in template • merge → PR publication → Scribe
returns-headingSOURCE / RETURN
outcomes-headingOUTCOMES
footerPR action can skip / fail and still reach Scribe. No-changes also reaches Scribe.
Agent workAgent
Agent workAgent task
Agent workagent
RAI gateRai
RAI gateVerdict routing
RAI gaterai
Human reviewReview
Human reviewhuman-review
MergeMerge
MergeMerge outcome routing
Mergemerge
Publish / reuse PRPublish / reuse PR
Publish / reuse PRCreate / reuse; not git push
Publish / reuse PRaction
ScribeScribe
ScribeRecord the run outcome
Scribescribe
Safety failedSafety failed
Safety failedWorkflow endpoint
DeclinedDeclined
DoneDone
edge-02-labelrevise
edge-03-labelsafety- failed
edge-04-labelno- changes
edge-05-labelreview
edge-06-labelapproved
edge-07-labelrequest-changes
edge-08-labeldeclined
edge-09-labelmerged
edge-10-labelblocked
Diagram details and constraints
ElementContract
titlePostgres is the event relay
subtitleAny API replica can serve a cursor over durable RunEvents—no sticky session required.
group-title0Write path · replica A
group-title1Read path · replica B
Run producerRun producer
Run producerAppend a structured event
Run producerrunId + type + payload
EF event streamEF event stream
EF event streamSerialize writes per run
EF event streampg_advisory_xact_lock
RunEventsRunEvents
RunEventsShared PostgreSQL table
RunEvents(RunId, Sequence)
Web / MCP watcherWeb / MCP watcher
Web / MCP watcherConsume ordered events
Web / MCP watcherlast delivered cursor
SSE endpointSSE endpoint
SSE endpointEmit id + event + data
SSE endpointordered response frames
EF subscriberEF subscriber
EF subscriberRead Sequence > cursor
EF subscriberidle poll: 250 ms
e1append
e2commit
e3ordered batch
e4yield
e5SSE frames
assurance-titlePOSTGRES LANE ONLY
assurance-line1SQLite register-channel / replay / tail is a separate implementation—not this architecture.
assurance-line2Late-delta suppression is process-local; do not read it as a database-wide terminal fence.
Run producerInput
Run producerRunStreamEntry
Run producerIdentity
Run producerrunId + event type
Run producerBody
Run producerStructured payload
Run producerAck
Run producerAfter durable commit
EF event streamLock
EF event streamPer-run advisory lock
EF event streamNext
EF event streamMAX(Sequence) + 1
EF event streamWrite
EF event streamSave transaction
EF event streamCommit
EF event streamBefore acknowledgement
RunEventsTable
RunEventsKey
RunEventsRunId + Sequence
RunEventsOrder
RunEventsAscending sequence
RunEventsReuse
RunEventsSame type / payload
Web / MCP watcherClient
Web / MCP watcherWeb or MCP
Web / MCP watcherResume
Web / MCP watcherLast delivered cursor
Web / MCP watcherReplica
Web / MCP watcherNo sticky requirement
Web / MCP watcherHistory
Web / MCP watcherDurable ordered events
SSE endpointFrame
SSE endpointid + event + data
SSE endpointCursor
SSE endpointLast-Event-ID
SSE endpointDelivery
SSE endpointYield ordered events
SSE endpointClose
SSE endpointAfter batch is drained
EF subscriberQuery
EF subscriberSequence > cursor
EF subscriberIdle
EF subscriberPoll after 250 ms
EF subscriberState
EF subscriberShared durable table
EF subscriberBlocked
EF subscriberRetryable: keep open
producerCoordinator or run execution; Acknowledgement follows commit
appendAllocate MAX(Sequence) + 1; Save and commit transaction
storeCross-replica ordered history; Explicit duplicates must match payload
clientReconnect from the cursor; No local channel dependency
sseCursor advances after delivery; Drain batch before terminal close
readerQuery the shared durable table; Retryable assembly_blocked stays open
notesPOSTGRES LANE ONLY; SQLite register-channel / replay / tail is a separate implementation—not this architecture.; Late-delta suppression is process-local; do not read it as a database-wide terminal fence.
groupsWrite path · replica A; Read path · replica B
Diagram details and constraints
ElementContract
titleOne heartbeat tick, two scopes
takeawayPickup runs per project; reconciliation, deferred-spec drain and optional reaping run afterward.
group-title-0PROJECT LOOP
group-title-1PER-PROJECT PICKUP
group-title-2ONCE AFTER THE PROJECT LOOP
Heartbeat tickHeartbeat tick
Heartbeat tickEnumerate projects
Heartbeat tickfailure-isolated sweep
Active + available?Active + available?
Active + available?Skip unavailable projects
Active + available?per-project admission
Capped Ready listCapped Ready list
Capped Ready listDeterministic candidates
Capped Ready listper-project limit
Atomic claimAtomic claim
Atomic claimReserve coordinator run
Atomic claimcompeting claim may lose
Start reserved runStart reserved run
Start reserved runCarry backlog origin
Start reserved runconfirmation policy applies
End project loopEnd project loop
End project loopRecord tick result
End project loopnot an inner-loop sweep
Reconcile onceReconcile once
Reconcile onceRepair durable supervision
Reconcile onceafter all projects
Drain spec decisionsDrain spec decisions
Drain spec decisionsRecover orphaned decisions
Drain spec decisionsdurable OutcomeSpec
Optional pod reaperOptional pod reaper
Optional pod reaperEvery N ticks when enabled
Optional pod reaperthrottled cleanup
e0each
e1eligible
e2claim
e3won
e4loop done
e5once
e6then
e7when due
groupsPROJECT LOOP; PER-PROJECT PICKUP; ONCE AFTER THE PROJECT LOOP
Diagram details and constraints
ElementContract
titleA dependency DAG, not agent chat
takeawayIllustrative tasks A-D show readiness; only assemble-ready/completed prerequisites count.
group-title-0DURABLE INTENT AND READINESS
group-title-1ILLUSTRATIVE PARALLEL ROOTS
group-title-2ILLUSTRATIVE DEPENDENTS AND HANDOFF
Confirmed OutcomeSpecConfirmed OutcomeSpec
Confirmed OutcomeSpecIntent before decomposition
Confirmed OutcomeSpecconfirmed status
Persisted WorkPlanPersisted WorkPlan
Persisted WorkPlanTasks + dependency edges
Persisted WorkPlanselected workflow
Readiness ruleReadiness rule
Readiness ruleEvery predecessor satisfied
Readiness ruleassemble_ready / done
Example root AExample root A
Example root ANo prerequisites
Example root Aillustrative, not fixed
Example root BExample root B
Example root Bparallel with A
Satisfied rootsSatisfied roots
Satisfied rootsNot merely terminal
Satisfied rootsfailure does not unlock
Example dependent CExample dependent C
Example dependent CDepends on A
Example dependent Cillustrative edge
Example dependent DExample dependent D
Example dependent DDepends on A and B
Example dependent Dillustrative join
Collective handoffCollective handoff
Collective handoffRecheck aggregate eligibility
Collective handoffquiescence != success
e0persist
e1evaluate
e2ready
e4A done
e6B done
e7both
e8settled
groupsDURABLE INTENT AND READINESS; ILLUSTRATIVE PARALLEL ROOTS; ILLUSTRATIVE DEPENDENTS AND HANDOFF
Diagram details and constraints
ElementContract
titleChild execution and dependency progress
takeawayChildren publish typed results and branch content; collective review belongs to the parent.
group-title-0DEPENDENCY ADMISSION
group-title-1TRIMMED CHILD EXECUTION
group-title-2PUBLISHED CONTENT AND AGGREGATE ELIGIBILITY
Pending subtasksPending subtasks
Pending subtasksPersisted dependency map
Pending subtasksnot automatically ready
Ready frontierReady frontier
Ready frontierAll predecessors satisfied
Ready frontiernot any terminal result
Isolated childIsolated child
Isolated childLaunch agent execution
Isolated childseparate checkout
Agent resultAgent result
Agent resultTyped conditional output
Agent resultno child RAI executor
Assemble-readyAssemble-ready
Assemble-readySuccessful child content
Assemble-readysatisfies dependents
Typed turn failureTyped turn failure
Typed turn failureDoes not satisfy dependents
Typed turn failureno per-child review
Published branchPublished branch
Published branchAuthoritative committed tip
Published branchnot shared mutable files
Dependency baseDependency base
Dependency baseRebuild prerequisite content
Dependency basenew isolated dependent
Parent assembly checkParent assembly check
Parent assembly checkQuiescence plus eligibility
Parent assembly checkcollective gates later
e0evaluate
e1dispatch
e2execute
e3success
e4failed
e5publish
e6integrate
e7unlock
e8settled
e9blocked
groupsDEPENDENCY ADMISSION; TRIMMED CHILD EXECUTION; PUBLISHED CONTENT AND AGGREGATE ELIGIBILITY
Diagram details and constraints
ElementContract
titlePersist the plan before dispatch
takeawayConfirmation, selection and decomposition precede durable child dispatch.
group-title-0INTENT AND REUSE
group-title-1SELECTION AND DURABLE PLAN
group-title-2DISPATCH AND COLLECTIVE HANDOFF
Submitted requestSubmitted request
Submitted requestGoal + caller context
Submitted requestmanual or pickup
Confirmation boundaryConfirmation boundary
Confirmation boundaryManual or unattended policy
Confirmation boundaryautopilot-dependent
Existing plan?Existing plan?
Existing plan?Reuse persisted plan
Existing plan?avoid decomposing twice
Select workflowSelect workflow
Select workflowAvailable definitions
Select workflowexplicit choices honored
Decompose + validateDecompose + validate
Decompose + validateOutcome-complete work
Decompose + validatecompatibility check
Persist WorkPlanPersist WorkPlan
Persist WorkPlanSubtasks and dependencies
Persist WorkPlanworkflow identity
Ready frontierReady frontier
Ready frontierDependency satisfaction
Ready frontierpending -> ready work
Dispatch childrenDispatch children
Dispatch childrenObserve classified outcomes
Dispatch childrenisolated child runs
Collective handoffCollective handoff
Collective handoffAfter child supervision
Collective handoffassembly eligibility
e0confirm
e1lookup
e2new
e3decompose
e4persist
e5reuse
e6ready
e7dispatch
e8handoff
groupsINTENT AND REUSE; SELECTION AND DURABLE PLAN; DISPATCH AND COLLECTIVE HANDOFF
Diagram details and constraints
ElementContract
titleExecution and observation cooperate
takeawayThe watcher projects runtime events into durable state; it is not the executing graph.
group-title-0START AND BIND
group-title-1RUNTIME AND SUPERVISION
group-title-2DURABLE PROJECTIONS
Run orchestratorRun orchestrator
Run orchestratorStarts factory + watcher
Run orchestratorsupervised lifetime
Effective definitionEffective definition
Effective definitionResolve concrete workflow
Effective definitionno policy composer
Factory + binderFactory + binder
Factory + binderBuild executable graph
Factory + bindertyped bindings
Checkpointed streamCheckpointed stream
Checkpointed streamMAF executes the graph
Checkpointed streamprovider-aware store
Watch loopWatch loop
Watch loopConsumes runtime updates
Watch loopnot graph execution
Review requestReview request
Review requestPersist pending decision
Review requestdurable pause context
Typed terminalTyped terminal
Typed terminalClassify completed output
Typed terminalnot inferred from text
Durable run stateDurable run state
Durable run statePersist status projection
Durable run statewatcher owns updates
Workflow-step eventsWorkflow-step events
Workflow-step eventsExpose execution progress
Workflow-step eventsclient observation
e0start
e1resolve
e2execute
e3supervise
e4stream
e5request
e6terminal
e7persist
e8publish
groupsSTART AND BIND; RUNTIME AND SUPERVISION; DURABLE PROJECTIONS
Diagram details and constraints
ElementContract
titleReview API decision paths
takeawayAuthorize first. Deliver through the right path. Lock before any merge CAS.
group-title-0ADMISSION AND REPLAY
group-title-1DELIVERY ALTERNATIVES
group-title-2CONTINUATION AND MERGE
Caller + accessCaller + access
Caller + accessProject contributor check
Caller + accesslegacy: pending owner
Reviewable state?Reviewable state?
Reviewable state?Inspect status + pending
Reviewable state?awaiting_review
Replay or conflictReplay or conflict
Replay or conflictMatching terminal: reuse
Replay or conflictotherwise: 409
Live pendingLive pending
Live pendingChanges / decline use CAS
Live pendingapprove: no merge CAS
Deferred pendingDeferred pending
Deferred pendingPersist the decision first
Deferred pendingthen status transition
No live / no pendingNo live / no pending
No live / no pendingValidate direct approval
No live / no pendingchanges: 409
Consume + deliverConsume + deliver
Consume + deliverSend workflow response
Consume + deliverlive continuation
Repository lockRepository lock
Repository lockOnly on reaching merge
Repository locklock before CAS
Merge CAS + GitMerge CAS + Git
Merge CAS + GitGuard reviewed tree input
Merge CAS + Gitrelease lock on exit
e0check
e1replay
e2live
e3deferred
e4direct
e5deliver
e6on merge
e7approve
e8locked
groupsADMISSION AND REPLAY; DELIVERY ALTERNATIVES; CONTINUATION AND MERGE