Distributed agents over A2A — Experience
Experimental transport
Distributed agent execution uses A2A and a pinned -preview package line. Sandbox:AgentExecutionMode selects in-process (in-api) or remote (pod-per-run) leaf turns. The code default is in-api; the checked-in Kubernetes API and worker deployments select pod-per-run. Changing this configuration is a deployment operation, not an instantaneous migration of active turns.
This doc describes remote execution for the person watching a run and the operator running the platform. The run and review model stays familiar, but pod placement, provisioning delays, and transport failures can be visible.
For the design, read the A2A bridge deep dive. For the surface and security gates, read the A2A reference. For the pod lifecycle itself, see Sandbox pod execution and its experience doc.
1. The headline: the run model stays in the worker
A remote leaf turn feeds the existing run timeline and workflow gates. From the user's seat:
- The worker records returned events in arrival order and exposes them through the existing event stream.
- Review and confirmation gates remain worker-owned; execution mode does not decide which gates a workflow requires.
- Coordinator launch mode still matters: Define Outcome requests confirmation, while Direct and unattended pickup do not add that manual gate.
There is no user-facing execution-mode switch. The topology can show a per-node pod indicator, and an interrupted remote turn can surface a structured failure. Familiar workflow semantics do not mean guaranteed gap-free streaming or invisible recovery.
2. What crosses the transport boundary
Only the leaf agent turn moves into a pod. The orchestration graph — including review gates and checkpoint management — stays in the worker. So:
- Workflow events remain worker-owned; the pod's turn events are decoded and recorded by the worker.
- The review/confirm gates are graph constructs that live in the worker; they never travel over the wire, so they behave exactly as before.
- The pod streams the turn's output back, and the worker re-injects it into the same event stream that feeds the browser.
The pod has no checkpoint-store or database connection. Durable checkpoints and run events are written through the worker, not directly by AgentHost. See the coordinator orchestration experience for the surrounding workflow.
3. What actually changes — and it is operational
The main operational differences are:
| Aspect | In-process (in-api) | Distributed (pod-per-run) |
|---|---|---|
| Where a turn runs | In the worker process | In a warm AgentHost sandbox pod configured for the run |
| Isolation | Shared worker process | Kata-isolated pod, scoped credential, default-deny egress |
| Memory footprint | Worker holds every active session | Heavy SDK session lives and dies in the pod |
| Failure blast radius | A bad turn can pressure the worker | A bad turn is contained to its pod |
| What an operator watches | Worker pods | Worker pods plus sandbox pods |
The operational wins are isolation and memory relief: the heavyweight model session leaves the worker process and runs in a disposable, isolated pod. A run no longer keeps a heavy session pinned in a shared process, which is the memory-pressure fix. And a misbehaving turn is contained inside its own Kata-isolated pod rather than sharing the worker's address space.
4. How to reason about it as an operator
A few mental models keep distributed execution easy to reason about.
The pod is disposable; recovery depends on durable state. Checkpoint state and the serialized session blob live outside the A2A connection. At a workflow review gate, RunWatchLoopService attempts pod release only when pod-per-run and Sandbox:ReleasePodOnSuspend=true are active. Resume uses recoverable checkpoint/session state and a newly claimed/configured pod. Release is best-effort, and retained assembly/preview resources are an exception: do not assume every human wait holds zero pods.
A dropped connection is not transparent replay. A2A's live stream has no mid-stream replay. Missing agent.turn.end produces retryable agent_host_turn_incomplete; transport exceptions become a2a_transport_failure, with retryability determined from the failure. These visible failures let the coordinator deliberately redispatch or follow its recovery policy. They do not guarantee uninterrupted output, duplicate-free side effects, or recovery without an operator decision.
The transport is the sole leaf-turn wire. Sandbox:AgentExecutionMode=in-api selects local leaf execution instead of another remote protocol. Apply the configuration through the deployment process and check running work and capacity; it does not move an already executing remote turn into the worker.
Every turn is authenticated to that run's pod. The worker path is API/worker → claim warm pod → POST /configure → RemoteAgentProxy → Authorization: Bearer {per-run token} → AgentHost message:stream. The token is generated at run launch, delivered by /configure, and accepted only by that pod. NetworkPolicy and mTLS still restrict who can reach the listener, but the turn endpoint also has application-layer bearer auth.
More pods to watch, same run model. The new operational surface is sandbox pods alongside worker pods. Their warm-pool sizing, isolation, and credential model are covered in sandbox pods reference. The run timeline, review gates, and event stream you already know are unchanged.
Check the actual transport configuration: the base AgentHost ConfigMap has RequireMtls=false for its PoC path, while the production overlay enables mTLS. Do not infer production TLS posture merely from the presence of a pod or a bearer token.
5. Where you see it: Web UI, MCP, and diagnostics
- Web UI. The existing timeline and review/merge surfaces remain. Pod indicators show recorded placement, and provisioning or remote-turn failures can appear in run state.
- MCP. The MCP tool surface that drives and observes runs is unchanged — starting, confirming, watching, steering, and reviewing a run work identically whether turns are in-process or distributed. (The MCP server itself is a separate inbound surface; see the MCP server deep dive and MCP client experience.)
- Diagnostics. Cluster health, warm-pool readiness, and per-run claims complement the run view. The checked-in AgentHost warm-pool target is two, not a guarantee that two pods are currently ready. East-west connectivity and bridge health belong to agent communication.
6. The one caveat to keep in mind
The transport dependency is preview-staged even though the checked-in Kubernetes deployments select remote execution. Treat code defaults, deployed configuration, and live cluster health as separate facts. The A2A reference and A2A bridge deep dive cover the transport contract; neither a preview dependency nor a mode switch promises transparent recovery.
Diagram details and constraints
| Element | Contract |
|---|---|
| title | Remote leaves, worker-owned graph |
| takeaway | A2A moves leaf turns into sandbox pods, not orchestration or durable state ownership. |
| group-title-0 | WORKER CONTROL PLANE |
| group-title-1 | LEAF EXECUTION AND DURABILITY |
| Workflow graph | Workflow graph |
| Workflow graph | Orchestration and gates |
| Workflow graph | RequestPort / checkpoints |
| Workflow graph | Graph progression and human gates stay worker-side. |
| RemoteAgentProxy | RemoteAgentProxy |
| RemoteAgentProxy | Leaf-turn adapter |
| RemoteAgentProxy | message:stream |
| RemoteAgentProxy | Configure run context; forward the leaf invocation. |
| Event recorder | Event recorder |
| Event recorder | Decode returned events |
| Event recorder | ordered sequence numbers |
| Event recorder | Records pod event data parts; no direct pod-to-UI stream. |
| Sandbox pod | Sandbox pod |
| Sandbox pod | AgentHost + executor |
| Sandbox pod | leaf agent execution |
| Sandbox pod | No database/checkpoint-store access from AgentHost. |
| Durable state | Durable state |
| Durable state | Checkpoints and events |
| Durable state | shared run state |
| Durable state | Worker persists progress; API reads event cursors. |
| Web timeline | Web timeline |
| Web timeline | API stream consumer |
| Web timeline | snapshot + SSE |
| Web timeline | Shows persisted run events; not transport-level replay. |
| e0 | invoke leaf |
| e1 | A2A call |
| e2 | RunEvents |
| e3 | record |
| e4 | API / SSE |
| note | Worker checkpoints and review gates remain authoritative. TLS settings depend on deployment overlay. |
| n0 | Graph progression and human gates stay worker-side. |
| n1 | Configure run context; forward the leaf invocation. |
| n2 | Records pod event data parts; no direct pod-to-UI stream. |
| n3 | No database/checkpoint-store access from AgentHost. |
| n4 | Worker persists progress; API reads event cursors. |
| n5 | Shows persisted run events; not transport-level replay. |
| groups | WORKER CONTROL PLANE; LEAF EXECUTION AND DURABILITY |
Diagram details and constraints
| Element | Contract |
|---|---|
| title | Remote turns: pause is not failure |
| takeaway | Claim/configure, observe terminal evidence, and make recovery an explicit policy decision. |
| group-title-0 | ACQUIRE AND EXECUTE |
| group-title-1 | DISTINCT OUTCOMES |
| Claim a sandbox | Claim a sandbox |
| Claim a sandbox | Controller/Kubernetes bind |
| Claim a sandbox | pod-per-run deployment |
| Claim a sandbox | Acquire a run-bound pod; not a pod for each token. |
| Configure + stream | Configure + stream |
| Configure + stream | Inject run context |
| Configure + stream | A2A message:stream |
| Configure + stream | Leaf output returns through the worker event recorder. |
| Terminal evidence | Terminal evidence |
| Terminal evidence | agent.turn.end |
| Terminal evidence | successful completion |
| Terminal evidence | A clean EOF without the terminal marker is not success. |
| Checkpoint wait | Checkpoint wait |
| Checkpoint wait | Human / external gate |
| Checkpoint wait | release is conditional |
| Checkpoint wait | Active previews or assembly may retain pod resources. |
| Visible failure | Visible failure |
| Visible failure | Incomplete / transport |
| Visible failure | retryable when classified |
| Visible failure | Prior deltas can be preserved; no seamless replay promise. |
| Recovery decision | Recovery decision |
| Recovery decision | Inspect state and budget |
| Recovery decision | redispatch when chosen |
| Recovery decision | Fresh dispatch is deliberate; side effects may need review. |
| e0 | configure |
| e1 | turn end |
| e2 | failure |
| e3 | evaluate |
| e4 | resume |
| note | The checkpoint lane is a separate workflow wait, not an automatic recovery path for a failed turn. |
| n0 | Acquire a run-bound pod; not a pod for each token. |
| n1 | Leaf output returns through the worker event recorder. |
| n2 | A clean EOF without the terminal marker is not success. |
| n3 | Active previews or assembly may retain pod resources. |
| n4 | Prior deltas can be preserved; no seamless replay promise. |
| n5 | Fresh dispatch is deliberate; side effects may need review. |
| groups | ACQUIRE AND EXECUTE; DISTINCT OUTCOMES |
