Skip to content

Distributed agents over A2A — Experience ​

Experimental transport

Distributed agent execution uses A2A and a pinned -preview package line. Sandbox:AgentExecutionMode selects in-process (in-api) or remote (pod-per-run) leaf turns. The code default is in-api; the checked-in Kubernetes API and worker deployments select pod-per-run. Changing this configuration is a deployment operation, not an instantaneous migration of active turns.

This doc describes remote execution for the person watching a run and the operator running the platform. The run and review model stays familiar, but pod placement, provisioning delays, and transport failures can be visible.

For the design, read the A2A bridge deep dive. For the surface and security gates, read the A2A reference. For the pod lifecycle itself, see Sandbox pod execution and its experience doc.

1. The headline: the run model stays in the worker ​

A remote leaf turn feeds the existing run timeline and workflow gates. From the user's seat:

  • The worker records returned events in arrival order and exposes them through the existing event stream.
  • Review and confirmation gates remain worker-owned; execution mode does not decide which gates a workflow requires.
  • Coordinator launch mode still matters: Define Outcome requests confirmation, while Direct and unattended pickup do not add that manual gate.

There is no user-facing execution-mode switch. The topology can show a per-node pod indicator, and an interrupted remote turn can surface a structured failure. Familiar workflow semantics do not mean guaranteed gap-free streaming or invisible recovery.

2. What crosses the transport boundary ​

Only the leaf agent turn moves into a pod. The orchestration graph — including review gates and checkpoint management — stays in the worker. So:

  • Workflow events remain worker-owned; the pod's turn events are decoded and recorded by the worker.
  • The review/confirm gates are graph constructs that live in the worker; they never travel over the wire, so they behave exactly as before.
  • The pod streams the turn's output back, and the worker re-injects it into the same event stream that feeds the browser.

The pod has no checkpoint-store or database connection. Durable checkpoints and run events are written through the worker, not directly by AgentHost. See the coordinator orchestration experience for the surrounding workflow.

3. What actually changes — and it is operational ​

The main operational differences are:

AspectIn-process (in-api)Distributed (pod-per-run)
Where a turn runsIn the worker processIn a warm AgentHost sandbox pod configured for the run
IsolationShared worker processKata-isolated pod, scoped credential, default-deny egress
Memory footprintWorker holds every active sessionHeavy SDK session lives and dies in the pod
Failure blast radiusA bad turn can pressure the workerA bad turn is contained to its pod
What an operator watchesWorker podsWorker pods plus sandbox pods

The operational wins are isolation and memory relief: the heavyweight model session leaves the worker process and runs in a disposable, isolated pod. A run no longer keeps a heavy session pinned in a shared process, which is the memory-pressure fix. And a misbehaving turn is contained inside its own Kata-isolated pod rather than sharing the worker's address space.

4. How to reason about it as an operator ​

A few mental models keep distributed execution easy to reason about.

The pod is disposable; recovery depends on durable state. Checkpoint state and the serialized session blob live outside the A2A connection. At a workflow review gate, RunWatchLoopService attempts pod release only when pod-per-run and Sandbox:ReleasePodOnSuspend=true are active. Resume uses recoverable checkpoint/session state and a newly claimed/configured pod. Release is best-effort, and retained assembly/preview resources are an exception: do not assume every human wait holds zero pods.

A dropped connection is not transparent replay. A2A's live stream has no mid-stream replay. Missing agent.turn.end produces retryable agent_host_turn_incomplete; transport exceptions become a2a_transport_failure, with retryability determined from the failure. These visible failures let the coordinator deliberately redispatch or follow its recovery policy. They do not guarantee uninterrupted output, duplicate-free side effects, or recovery without an operator decision.

The transport is the sole leaf-turn wire. Sandbox:AgentExecutionMode=in-api selects local leaf execution instead of another remote protocol. Apply the configuration through the deployment process and check running work and capacity; it does not move an already executing remote turn into the worker.

Every turn is authenticated to that run's pod. The worker path is API/worker → claim warm pod → POST /configure → RemoteAgentProxy → Authorization: Bearer {per-run token} → AgentHost message:stream. The token is generated at run launch, delivered by /configure, and accepted only by that pod. NetworkPolicy and mTLS still restrict who can reach the listener, but the turn endpoint also has application-layer bearer auth.

More pods to watch, same run model. The new operational surface is sandbox pods alongside worker pods. Their warm-pool sizing, isolation, and credential model are covered in sandbox pods reference. The run timeline, review gates, and event stream you already know are unchanged.

Check the actual transport configuration: the base AgentHost ConfigMap has RequireMtls=false for its PoC path, while the production overlay enables mTLS. Do not infer production TLS posture merely from the presence of a pod or a bearer token.

5. Where you see it: Web UI, MCP, and diagnostics ​

  • Web UI. The existing timeline and review/merge surfaces remain. Pod indicators show recorded placement, and provisioning or remote-turn failures can appear in run state.
  • MCP. The MCP tool surface that drives and observes runs is unchanged — starting, confirming, watching, steering, and reviewing a run work identically whether turns are in-process or distributed. (The MCP server itself is a separate inbound surface; see the MCP server deep dive and MCP client experience.)
  • Diagnostics. Cluster health, warm-pool readiness, and per-run claims complement the run view. The checked-in AgentHost warm-pool target is two, not a guarantee that two pods are currently ready. East-west connectivity and bridge health belong to agent communication.

6. The one caveat to keep in mind ​

The transport dependency is preview-staged even though the checked-in Kubernetes deployments select remote execution. Treat code defaults, deployed configuration, and live cluster health as separate facts. The A2A reference and A2A bridge deep dive cover the transport contract; neither a preview dependency nor a mode switch promises transparent recovery.

Diagram details and constraints
ElementContract
titleRemote leaves, worker-owned graph
takeawayA2A moves leaf turns into sandbox pods, not orchestration or durable state ownership.
group-title-0WORKER CONTROL PLANE
group-title-1LEAF EXECUTION AND DURABILITY
Workflow graphWorkflow graph
Workflow graphOrchestration and gates
Workflow graphRequestPort / checkpoints
Workflow graphGraph progression and human gates stay worker-side.
RemoteAgentProxyRemoteAgentProxy
RemoteAgentProxyLeaf-turn adapter
RemoteAgentProxymessage:stream
RemoteAgentProxyConfigure run context; forward the leaf invocation.
Event recorderEvent recorder
Event recorderDecode returned events
Event recorderordered sequence numbers
Event recorderRecords pod event data parts; no direct pod-to-UI stream.
Sandbox podSandbox pod
Sandbox podAgentHost + executor
Sandbox podleaf agent execution
Sandbox podNo database/checkpoint-store access from AgentHost.
Durable stateDurable state
Durable stateCheckpoints and events
Durable stateshared run state
Durable stateWorker persists progress; API reads event cursors.
Web timelineWeb timeline
Web timelineAPI stream consumer
Web timelinesnapshot + SSE
Web timelineShows persisted run events; not transport-level replay.
e0invoke leaf
e1A2A call
e2RunEvents
e3record
e4API / SSE
noteWorker checkpoints and review gates remain authoritative. TLS settings depend on deployment overlay.
n0Graph progression and human gates stay worker-side.
n1Configure run context; forward the leaf invocation.
n2Records pod event data parts; no direct pod-to-UI stream.
n3No database/checkpoint-store access from AgentHost.
n4Worker persists progress; API reads event cursors.
n5Shows persisted run events; not transport-level replay.
groupsWORKER CONTROL PLANE; LEAF EXECUTION AND DURABILITY
Diagram details and constraints
ElementContract
titleRemote turns: pause is not failure
takeawayClaim/configure, observe terminal evidence, and make recovery an explicit policy decision.
group-title-0ACQUIRE AND EXECUTE
group-title-1DISTINCT OUTCOMES
Claim a sandboxClaim a sandbox
Claim a sandboxController/Kubernetes bind
Claim a sandboxpod-per-run deployment
Claim a sandboxAcquire a run-bound pod; not a pod for each token.
Configure + streamConfigure + stream
Configure + streamInject run context
Configure + streamA2A message:stream
Configure + streamLeaf output returns through the worker event recorder.
Terminal evidenceTerminal evidence
Terminal evidenceagent.turn.end
Terminal evidencesuccessful completion
Terminal evidenceA clean EOF without the terminal marker is not success.
Checkpoint waitCheckpoint wait
Checkpoint waitHuman / external gate
Checkpoint waitrelease is conditional
Checkpoint waitActive previews or assembly may retain pod resources.
Visible failureVisible failure
Visible failureIncomplete / transport
Visible failureretryable when classified
Visible failurePrior deltas can be preserved; no seamless replay promise.
Recovery decisionRecovery decision
Recovery decisionInspect state and budget
Recovery decisionredispatch when chosen
Recovery decisionFresh dispatch is deliberate; side effects may need review.
e0configure
e1turn end
e2failure
e3evaluate
e4resume
noteThe checkpoint lane is a separate workflow wait, not an automatic recovery path for a failed turn.
n0Acquire a run-bound pod; not a pod for each token.
n1Leaf output returns through the worker event recorder.
n2A clean EOF without the terminal marker is not success.
n3Active previews or assembly may retain pod resources.
n4Prior deltas can be preserved; no seamless replay promise.
n5Fresh dispatch is deliberate; side effects may need review.
groupsACQUIRE AND EXECUTE; DISTINCT OUTCOMES