Skip to content

Microsoft Agent Framework — Conceptual Deep Dive ​

Purpose & Mental Model ​

Agentweaver uses the Microsoft Agent Framework (MAF), shipped as Microsoft.Agents.AI.Workflows (with checkpointing), for full/child workflow graphs and the coordinator's spec/confirmation phase. MAF schedules typed executors, checkpoints, suspends, and resumes those graphs. Service-driven dispatch/assembly and Operator conversations are deliberately different execution paths.

The useful boundary is: workflow-bound work uses a MAF graph, agent execution is a leaf, and MAF human gates use RequestPort. Collective assembly approval instead uses AssemblyReviewGate and durable service recovery. Operator conversations use their own durable event history rather than a review/merge graph.

This deep dive answers two questions: what MAF gives Agentweaver, and why Agentweaver builds on it instead of a bespoke engine. For the wider picture, see the system overview.

Why not a bespoke engine? Three properties are hard to get right and MAF supplies them as primitives:

  • Durable suspend/resume. A run that pauses for a human must survive a process restart and pick up exactly where it stopped. MAF makes this a first-class operation (checkpoint + resume from a request port).
  • Typed graph composition with adapters. Nodes have typed inputs and outputs; edges connect them; cross-type transitions are explicit adapters. This makes "fail closed when a node can't run" a structural property, not a runtime hope.
  • A standard agent leaf. The AIAgent abstraction is the unit MAF schedules and the unit A2A remotes — so the same abstraction carries both in-process and distributed execution.

Building those three from scratch is the expensive part of an orchestration runtime. Agentweaver spends its effort on the policy (which graph, which gates, which agents) and lets MAF own the mechanism.

The Workflow graph: typed executors and edges ​

A MAF workflow is a directed graph. Each node is an executor with a typed input and a typed output. Each edge carries a value of the producer's output type into a consumer that accepts it; when the types don't line up, an adapter executor sits on the edge to transform one into the other. MAF runs the graph in supersteps, delivering each executor its input, collecting its output, and routing along the edges whose (optional) predicate matches.

Agentweaver assembles a run's graph from a WorkflowDefinition (the declarative policy graph described in workflow-engine.md). Binding turns each logical node into its own MAF executor:

  • agent-turn node → an AgentTurnExecutor wrapping one AIAgent. This is the leaf unit of work.
  • peer-review node → an AI reviewer executor that emits an approve / request-changes verdict.
  • merge node → a MergeExecutor that applies the produced tree.
  • scribe node → a recording executor that captures the outcome.
  • rai / human-review gates → policy gates and request ports.

The binder mints a distinct executor per logical node, keyed by node id. This is why chained turns each get their own node: a workflow with three sequential agent turns produces three separate MAF executors, not one executor invoked three times. Distinct nodes are what make the topology graph legible and what let MAF emit a clean lifecycle event per step. Edges that cross types — AgentTurnOutput into a review request, a review decision into a merge input — are expanded into adapter executors so the typed contract is never violated. Binding fails closed: a node kind or edge with no executor mapping aborts the build instead of becoming a silent no-op.

The AIAgent abstraction and CopilotAIAgent ​

The leaf unit of work in a MAF graph is an AIAgent. Hosted execution uses a remote AgentHost proxy at that seam; the pod-side agent wraps the SDK session. The graph remains in the orchestration process. CopilotAIAgent is the SDK-backed implementation, not a separate deployed orchestration service.

The decisive property is that the MAF checkpoint manager can serialize the agent. CopilotAIAgent exposes its Copilot SDK session as a serializable blob, so when MAF writes a checkpoint it persists the agent's session into that checkpoint alongside the workflow's superstep state. This is what makes a paused run truly durable: resuming restores not just "which node we stopped at" but the agent's own conversational state. A bespoke engine would have to invent and test this session-serialization contract; MAF makes the agent a serializable graph citizen for free.

A fresh worker agent is minted per run through an injectable factory seam (IWorkflowAgentFactory), so production builds a CopilotAIAgent while tests substitute a fake — without changing the graph. Auxiliary turns (RAI, scribe, peer-review) each construct their own ephemeral agents the same way.

Executor lifecycle → live UI ​

MAF emits a typed event for every executor's lifecycle: ExecutorInvokedEvent when a node starts, ExecutorCompletedEvent when it finishes, and ExecutorFailedEvent when it throws. It also emits RequestInfoEvent when the graph reaches a request port and WorkflowOutputEvent at terminal output. These are the same WorkflowEvent stream MAF uses to drive execution.

The run watch loop subscribes to this stream and translates each lifecycle event into a workflow.step event on the run's own stream: invoked → started, completed → completed, failed → failed. Those workflow.step events are what drive the live topology graph in the UI — each node lights up as MAF invokes and completes its executor. The watch loop deliberately skips nodes that emit their own richer events (agent, rai, merge, scribe, review), so the timeline never double-reports. The result is that the browser's animated graph is a faithful projection of MAF's real scheduling, not a separately maintained model.

RequestPort: the MAF human-gate seam ​

MAF-bound human gates use a RequestPort with typed request and response. Routing a value into the port emits RequestInfoEvent, enters PendingRequests, and suspends that graph until the matching response arrives. This does not describe collective assembly approval or Assistant tool approvals.

This seam implements the following graph gates:

  • The run review gate is RequestPort.Create<WorkflowReviewRequest, WorkflowReviewDecision>("review-gate"). The agent's output is adapted into a WorkflowReviewRequest, the run suspends, and the reviewer's approve / request-changes / decline becomes a WorkflowReviewDecision.
  • The coordinator's OutcomeSpec confirmation gate is RequestPort.Create<CoordinatorOutcomeSpecRequest, CoordinatorOutcomeSpecDecision>. The drafted spec suspends the coordinator run until a human confirms or revises.
  • Per-node human-review gates in catalog/generated workflows are minted the same way, one request port per human-review node.

Because all of these are the same primitive, the suspend/resume plumbing is written once. When the watch loop sees a RequestInfoEvent it records the pending request, marks the run awaiting review, and closes the live stream at the gate. When the human responds, the decision is persisted against the exact request id and decision identity, claimed for delivery, and sent back into the suspended workflow. Review delivery is marked complete when its watch loop observes workflow progress; coordinator OutcomeSpec delivery is marked complete when SendResponseAsync acknowledges the correlated response, before unrelated stream events can acknowledge it accidentally. A crash before send leaves the decision retryable; a delivered decision no-ops during recovery instead of advancing the gate twice. The merge-blocked retry path even re-enters the same review gate, so a transient block keeps the workflow alive instead of failing it.

Checkpointing & durable resume ​

MAF persists workflow state through a checkpoint store — any implementation of ICheckpointStore<JsonElement>, which MAF exposes as the JsonCheckpointStore base class. Agentweaver selects the store at startup based on Database:Provider via ICheckpointStoreFactory (apps/Agentweaver.Api/Infrastructure/ICheckpointStoreFactory.cs):

  • Production (Postgres) → a shared, concurrency-safe PostgresJsonCheckpointStore. This is the default on hosted deployments and the correct fix for multi-replica operation. It derives from MAF's JsonCheckpointStore (so it plugs straight into CheckpointManager.CreateJson(store)) and persists every checkpoint as one row in the workflow_checkpoints table (apps/Agentweaver.Api/Infrastructure/Ef/PostgresJsonCheckpointStore.cs). Because each checkpoint is an independent, unique-PK row, the two API replicas write concurrently as plain INSERTs that never contend — there is no exclusive lock — and Postgres MVCC makes a committed checkpoint immediately visible to the other replica. That is genuine cross-pod checkpoint sharing and resume: a run suspended on pod A can be resumed from pod B.
  • Local / dev (SQLite or no database) → the file store. MAF's FileSystemJsonCheckpointStore, wrapped by ResilientCheckpointStore (apps/Agentweaver.Api/Infrastructure/ResilientCheckpointStore.cs), which hardens single-node startup so the API always boots. This path is no longer the production default; it remains for the single-writer dev experience and as a defensive safety net.

Why Postgres — the file store cannot be shared across replicas ​

FileSystemJsonCheckpointStore takes an exclusive process lock on its directory. The API runs replicas: 2 with HOME on a shared RWX Azure Files volume, so only one pod could ever hold that lock; the other was forced to a per-pod directory and the two replicas never shared checkpoints — cross-replica resume was impossible no matter how the volume permissions were set. Quieting the resulting log noise or fixing permissions only treated symptoms; the architectural fix is to move checkpoints into the database the app already runs, where concurrent writers are a first-class operation. The workflow_checkpoints schema:

ColumnPurpose
store_nameDiscriminator partitioning the two logical stores that were previously separate directories: runs (RunWorkflowFactory) and coordinator (CoordinatorWorkflowFactory).
session_idMAF session id — the RunId for the runs store.
checkpoint_idUnique GUID generated on create.
parent_checkpoint_id, has_parent_metadataMirror MAF's FileSystem index semantics so the parent-scoped index query behaves identically.
payload (jsonb)The checkpoint document.
created_at, updated_atTimestamps.

Primary key (store_name, session_id, checkpoint_id); index on (store_name, session_id). Concurrency is guaranteed structurally: fresh-GUID checkpoint ids mean every create is a non-conflicting INSERT, so two replicas never collide and no global lock is needed. The store uses IDbContextFactory<MemoryDbContext> (a fresh context per call), so a single registered instance serves many concurrent runs. The migration is Agentweaver.Api.Migrations.Postgres/Migrations/20260628140000_AddWorkflowCheckpoints.cs, applied on startup by the same MemoryDbContext.MigrateAsync() as every other table.

The file store's startup safety net (dev / fallback only) ​

When the file store is in use, ResilientCheckpointStore still guards three single-node hazards so the API never crash-loops:

  • Corrupt index. MAF parses index.jsonl one JSON object per line at construction, so a blank or partially-written line throws. The factory sanitizes the index (dropping unparseable lines after backing up the original) and quarantines an unrecoverable index instead of crash-looping. Genuine corruption is logged loudly (error) and quarantined; quarantine destinations are unique per pod and per call (index.jsonl.corrupt.{podId}.{unixSeconds}.{guid}, moved with overwrite) so rapid restarts cannot collide with IOException: already exists.
  • Multi-writer lock contention and shared-volume permission denial. If two processes share the directory or the volume is not writable, this is not corruption: ResilientCheckpointStore recognises it (walking the inner-exception chain for IsAccessDenied), skips quarantine, and falls back quietly (at most one concise warn per store, no fail/stacktrace) to a per-pod sub-directory so the node still boots. On Postgres these cases simply do not arise, because there is no shared file and no exclusive lock.

The selected store is handed to a CheckpointManager, and the manager checkpoints around every suspension.

Cross-replica resume is now real on Postgres

On the production Postgres path, both replicas read and write the same workflow_checkpoints rows, so a run suspended on one pod resumes on the other. The previous per-pod file fallback (where each replica checkpointed to its own directory and cross-replica resume was impossible) applies only to the SQLite/dev file store.

Two facts make resume robust:

  • The runId is the MAF session id. A run's id is used directly as MAF's session identifier, so a run's checkpoints are keyed by that id (the session_id column on Postgres, or a directory on the file store) and the most recent checkpoint is the resume point. There is no separate mapping table to keep consistent.
  • A checkpoint carries both the superstep state and the serialized agent session, including the correlation id of any suspended request port. Restoring a checkpoint rehydrates the graph and the agent, then continues from the gate.

On process restart, the WorkflowRestartService reconciles interrupted runs. A run recorded as awaiting review is resumed from its latest checkpoint: it rebuilds the workflow shape, calls MAF's resume-from-checkpoint, and restarts the watch loop so the run lands back at its suspended gate. If no checkpoint exists, a stale review is failed closed, while a still-valid one re-emits a synthetic review.requested after revalidating the worktree. Coordinator runs still in their spec phase are recovered the same way through the CoordinatorWorkflowFactory, which resumes the suspended confirmation gate from its own checkpoint.

Ordinary response versus process restoration ​

A live workflow receives the correlated decision through SendResponseAsync. A request arriving on another replica persists a ready delivery record for the owning watch loop instead of deleting the pending gate first. The delivery record moves waiting -> ready -> delivering -> delivered; stale delivering claims are retried, while delivered records no-op. This covers both the run review gate and the coordinator OutcomeSpec confirmation gate; the remaining legacy request-changes endpoint may destructively remove only a still-waiting gate because that human action abandons the paused workflow and starts a fresh revision rather than resuming automated parent work. This is a durable resume-delivery fence for Agentweaver-controlled review/resume handoffs, not a claim of exactly-once external side effects. If arbitrary shell, network, or Git effects happened after a resume, those phases still need their own retry or probe-and-reconcile semantics.

Process-loss recovery instead loads the selected checkpoint store, rebuilds the appropriate full/child graph, calls ResumeStreamingAsync, and restarts observation. Recovery revalidates durable state and worktree/tree identity; it is not unconditional success.

On multi-replica deployments, startup recovery first claims the run's durable execution lease before classifying an in_progress run as abandoned. A run with an unexpired peer lease is left untouched. If a running watch loop later loses its fencing token during lease renewal, that superseded owner stops without publishing a terminal transition; the new owner is responsible for recovery.

Postgres checkpoints are shared rows. File checkpoints are the SQLite/dev provider choice, not an automatic fallback when production Postgres fails.

Where Agentweaver deliberately does NOT use MAF (decision D3) ​

MAF is the right tool for a graph that pauses for humans and must survive restarts. It is not the right tool for everything, and Agentweaver draws a deliberate boundary.

The coordinator's spec/confirm phase is a MAF workflow: draft → RequestPort confirmation gate → confirm-terminal | revise-loop. That phase needs exactly what MAF provides — it drafts an outcome spec, suspends on a human confirmation gate, and must resume that gate after a restart. So it is checkpointed and resumable just like the run review gate.

After the human confirms the spec, the coordinator hands off to a service-driven engine — not a MAF graph (decision D3). Subtask dispatch and collective assembly run as background services whose entire state lives in database rows: the WorkPlan, the subtask DAG and its dependency edges, child run rows, and assembly status. The assembly pipeline reuses the real executors (RAI, scribe, merge plumbing) but invokes them directly, passing a NoOpWorkflowContext — a stub IWorkflowContext that throws on state operations — precisely to prove these calls do not depend on a live workflow graph.

The reasoning is the core of D3: MAF checkpoints exist to make in-memory graph state durable across suspension; the dispatch and assembly phases have no in-memory graph state worth checkpointing because their state is already durable in the DB. A coordinator can dispatch ten children, observe them, and assemble their branches entirely from persisted rows. If the process dies, recovery re-reads those rows and re-arms dispatch — no checkpoint required. Forcing those phases into a MAF graph would add a second source of truth (checkpoint and DB rows) that must be kept consistent, for no durability gain. So the boundary is: MAF where a run suspends on a human and resumes in-memory; service-driven where state is naturally relational and long-lived.

Collective Git finalization has its own durable effect protocol on the WorkPlan row. Before moving the originating branch, Agentweaver records an immutable effect id, lifecycle generation, repository identity, exact source ref/commit/tree, exact target ref and old commit, intended result commit/tree, and any checked-out-worktree pre-state. Under the repository merge lock it revalidates run ownership and lifecycle authorization, then advances the target with native git update-ref --no-deref <ref> <new> <old>. A restart proves applied only when the exact intended commit is at or in the target history with the recorded parent/precondition structure; it proves not_applied only while the target is still the exact old commit and the source was not already reachable. Same-tree commits, unrelated target movement, changed source refs, or unsafe checkout convergence are never treated as proof.

An ambiguous observation is persisted as assembly_unknown with evidence and an operator action, and is not automatically re-armed. A proven applied recovery finalizes without replaying the merge or Scribe. This gives the Git ref mutation a crash-safe, evidence-based recovery boundary; it does not claim general exactly-once semantics for shell commands, network calls, Scribe exports, or other external effects.

Each child run, however, is itself an ordinary MAF run with its own graph — so MAF still orchestrates every leaf of real work. D3 is only about the coordinator's dispatch/assembly tier, not the workers it launches.

A2A is also MAF ​

Distributed execution does not move the graph; it moves only the leaf. The worker↔pod transport, A2A, ships in the same .NET Agent Framework line and remotes at the AIAgent seam. The worker keeps the whole MAF graph — every executor, every WorkflowEvent, every RequestPort/HITL gate — and replaces only the leaf AIAgent with a proxy that forwards a single turn to a sandbox pod and streams the result back.

This is why the MAF-centric design here stays intact under distribution: no MAF event crosses the wire, no gate crosses the wire, and there is no MAF↔A2A translation layer, because only the leaf's streaming response travels. Checkpoints and the serialized session blob still live on the worker's checkpoint store, so durable resume is unchanged. The full reasoning — cut at the leaf not the graph, message-mode only, the RunEvent side-channel codec — is in the A2A bridge deep dive.

Invariants to preserve when rebuilding ​

  • A workflow-bound run uses a MAF graph of typed executors/edges; Operator conversations and service-driven collective phases do not.
  • A WorkflowDefinition binds to one MAF executor per logical node; chained turns get distinct nodes; binding fails closed.
  • The leaf is an AIAgent; production uses CopilotAIAgent, whose Copilot session the checkpoint manager serializes into the checkpoint.
  • MAF lifecycle events (ExecutorInvoked/Completed/Failed) are translated by the watch loop into workflow.step events that drive the live topology graph.
  • MAF human gates are RequestPorts; collective assembly uses a service gate backed by durable recovery instead.
  • Checkpoints use a provider-selected ICheckpointStore<JsonElement>: on Postgres the shared PostgresJsonCheckpointStore (rows in workflow_checkpoints, no lock, cross-replica resume), on SQLite/dev the FileSystemJsonCheckpointStore wrapped by ResilientCheckpointStore. The runId is the MAF session id; restart recovery resumes a suspended run from its latest checkpoint at the gate. On Postgres both replicas: 2 read/write the same rows; the old per-pod file fallback (where the losing replica took its own directory and cross-replica resume was impossible) applies only to the file store.
  • The coordinator's spec/confirm phase is MAF; dispatch and collective assembly (D3) are service-driven over DB rows, with no MAF graph and a NoOpWorkflowContext for direct executor calls.
  • A2A remotes only the AIAgent leaf; the MAF graph and all WorkflowEvent/RequestPort logic stay in the worker.
Diagram details and constraints
ElementContract
titleTyped MAF adapters · state completes the contract
takeawayReview decisions carry approval; saved AgentTurnOutput supplies the merge data.
group-0-titleREVIEW / STATE
group-1-titleRESPONSE / MERGE
AgentTurnOutputAgentTurnOutput
AgentTurnOutputRepository / worktree / tree data
AgentTurnOutputSuccessful turn provides the merge contract
AgentTurnOutputRunWorkflowFactory:409–440
Review adapterReview adapter
Review adapterSave output in workflow state
Review adapterEmit a WorkflowReviewRequest
Workflow stateWorkflow state
Workflow stateSaved AgentTurnOutput
Workflow stateDecision alone cannot reconstruct merge input
RequestPortRequestPort
RequestPortSuspend for external reviewer
RequestPortCorrelate request with WorkflowReviewDecision
Approved decisionApproved decision
Approved decisionWorkflowReviewDecision
Approved decisionApproval authorizes merge, not its payload
Merge adapterMerge adapter
Merge adapterRead the saved output
Merge adapterCombine approval + state into MergeInput
Blocked adapterBlocked adapter
Blocked adapterMergeOutput says blocked
Blocked adapterUse saved output to request another review
Blocked adapterRunWorkflowFactory:532–545
Merge executorMerge executor
Merge executorConsume typed MergeInput
Merge executorBlocked is retriable, not a completed merge
AgentTurnOutputturn output
Review adaptersave output
RequestPortdecision
Workflow stateread state
Approved decisionapproved
Merge adapterMergeInput
Merge executorblocked
scopeRepresentative full-run adapters only. Collective assembly and Operator conversations are different paths.
groupsREVIEW / STATE; RESPONSE / MERGE
Diagram details and constraints
ElementContract
titleCoordinator · MAF hands off to services
takeawayConfirmed spec can start dispatch; later collective phases are relational and service-driven.
group-0-titleSPEC HANDOFF
group-1-titleSERVICE-DRIVEN COLLECTIVE
CoordinatorOutcomeCoordinatorOutcome
CoordinatorOutcomeSpec confirmed
CoordinatorOutcomeCheck persisted work plan and subtasks
CoordinatorOutcomeCoordinatorRunService:1172–1250
StartDispatchStartDispatch
StartDispatchAuto-dispatch + nonempty plan
StartDispatchConfirmation alone is insufficient
Release MAF stateRelease MAF state
Release MAF stateRegistry and checkpoints released
Release MAF stateCoordinator run + event stream remain active
Dispatch / assemblyDispatch / assembly
Dispatch / assemblyService drivers own later phases
Dispatch / assemblyRead and update relational phase state
Work plan + child runsWork plan + child runs
Work plan + child runsRelational durable coordination
Work plan + child runsSubtasks retain artifact and phase information
Direct executor callsDirect executor calls
Direct executor callsRai / rubberduck / build-test / Scribe
Direct executor callsHandleAsync with NoOpWorkflowContext.Instance
Direct executor callsCollectiveAssemblyPipeline
Restart recoveryRestart recovery
Restart recoveryRead persisted coordinator phase
Restart recoveryRe-arm services; only spec resumes MAF
Assembly review gateAssembly review gate
Assembly review gateService-owned pending approval
Assembly review gateTaskCompletionSource, not MAF RequestPort
Assembly review gateAssemblyReviewGate:6–44
CoordinatorOutcomehandoff
CoordinatorOutcomerelease
StartDispatchstart
Dispatch / assemblypersist
Dispatch / assemblyinvoke
Work plan + child runspersisted
Direct executor callsservice
Restart recoveryre-arm
scopeNo checkpointed MAF assembly graph. Operator history/session handling is a separate exception.
groupsSPEC HANDOFF; SERVICE-DRIVEN COLLECTIVE
Diagram details and constraints
ElementContract
titleResponse delivery is not process restoration
takeawayLive response, another-replica delivery and process recovery are distinct control paths.
group-0-titleRESPONSE DELIVERY
group-1-titleCHECKPOINT / PROCESS RECOVERY
Reviewer / API checksReviewer / API checks
Reviewer / API checksRole, state and provider validation
Reviewer / API checksCorrelate the pending request identifier
Reviewer / API checksRunEndpoints:887–964
Existing StreamingRunExisting StreamingRun
Existing StreamingRunLocal workflow exists
Existing StreamingRunSendResponseAsync continues suspended port
Existing StreamingRunRunEndpoints:1052–1060
Durable deferred decisionDurable deferred decision
Durable deferred decisionNo local workflow, pending request
Durable deferred decisionPersist decision for the owning replica
Owning watch loopOwning watch loop
Owning watch loopPoll pending decisions
Owning watch loopDeliver via SendResponseAsync, not restore
Owning watch loopRunWatchLoopService:164–227
CheckpointManagerCheckpointManager
CheckpointManagerMAF saves workflow checkpoint
CheckpointManagerProvider selection: shared PG rows or local files
CheckpointManagerProgram.cs:1058–1065
Recovery serviceRecovery service
Recovery serviceRead latest checkpoint + run
Recovery serviceLoad persisted effective definition
Recovery serviceRunWorkflowFactory:1561–1586
Rebuild full / child graphRebuild full / child graph
Rebuild full / child graphChoose the correct persisted definition
Rebuild full / child graphMissing worktree/tree mismatches can fail closed
Rebuild full / child graphWorkflowRestartServiceTests
ResumeStreamingAsyncResumeStreamingAsync
ResumeStreamingAsyncRestore process state
ResumeStreamingAsyncRestart watch loop; recover suspended gate
Reviewer / API checkslocal
Reviewer / API checksother
Durable deferred decisionpoll decision
CheckpointManagercheckpoint
Recovery servicedefinition
Rebuild full / child graphresume
scopePostgreSQL failure does not select local files. Ordinary approval does not restore a checkpoint.
groupsRESPONSE DELIVERY; CHECKPOINT / PROCESS RECOVERY