Skip to content

Sandbox pod execution experience ​

This page is about what a person sees and feels when Agentweaver runs each agent in its own per-run sandbox pod — both the user watching a run in the web UI and the operator reasoning about the cluster. For the design logic see the Sandbox pod execution deep dive; for flags, identity, and the token mechanism see the Sandbox pods reference.

Related journeys: Runs, board & watch, Coordinator & orchestration, Operations, and the A2A distributed agents experience.

Mental model ​

The most important thing to feel about pod-per-run is that almost nothing changes in how a run looks. The board, live timeline, coordinator topology, and review gates retain their existing roles. What changes is where the work physically runs: each run's agent now executes inside its own Kata-isolated pod instead of inside the worker process. A pod name on the agent box shows recorded placement. Provisioning delays and remote-turn failures can also be visible.

A well-behaved run should make the user barely notice the sandbox. The pod pill is a placement signal, not independent proof of the sandbox's full isolation or credential posture.

What the pod pill is ​

When the backend runs inside Kubernetes, each agent box in a run's topology shows a small pod pill — a compact, monospace pill with a server icon and the executing pod's name (for example agentweaver-api-abc123 or agent-pod-worker-7). Hovering it shows the tooltip "Executing in pod {name}", and it is exposed to assistive technology with the same label.

What the user can take from it:

  • Where this agent is actually running. Under pod-per-run, the pill on a node is that run's sandbox pod, so two concurrently-running agents can show two different pod names — a direct, visible cue that their work is isolated from each other.
  • Which placement was recorded. Different pod names distinguish execution locations; configuration and cluster policy, not the pill alone, establish isolation.

The pill is Kubernetes-only and quiet by design:

  • on local/dev runs, or anywhere the backend is not in Kubernetes, no pill is shown — the UI stays clean and there is nothing to explain;
  • if the pod name is not yet known (the claim has not bound), no pill is shown rather than a placeholder; and
  • if the runtime probe fails, the UI degrades silently to no pill rather than erroring.

Where the value comes from: the coordinator graph passes each node's executionPodName directly to the pod indicator. A missing child-pod identity stays absent; it is not replaced with the API replica's name. The binding is persisted in the shared run-event log so another replica can serve the recorded placement. The data behind it is in the reference.

What happens during a run ​

From the user's side, submitting and watching a run feels identical to before:

  1. They submit work (or a coordinator goal) and watch the board / topology as usual.
  2. Agent boxes light up as running, stream tokens and status live, and — on Kubernetes — show their pod pill.
  3. The worker records pod-produced events and serves the existing timeline. Provisioning delays or an interrupted remote turn can still become visible.
  4. When the run completes, the box settles into its terminal state as usual.

Under the hood the heavy work — the model session, the tools, the shell — is happening in the pod, and the worker is relaying its stream into the same timeline. The user does not have to know that; the only new thing on screen is the pod name.

Sandbox preview: reaching a server inside the pod ​

Sometimes an agent starts a server inside its sandbox pod — a dev server, a freshly built app, a debug endpoint — and a person wants to actually open it and look. Because the pod is isolated, that server is not reachable by default. The sandbox browser preview publishes a run-scoped HTTPS capability URL through the preview Gateway; the API provisions the route, rather than proxying app traffic.

A Preview Sandbox button appears on the coordinator run page only when the run is using the Kubernetes sandbox (sandboxBackend === 'kubernetes-sandbox-claim', read from the run's sandbox.selected event) and a preview lifecycle state or existing preview session is present. The current button is not gated simply on whether the run is active; retained previews can be opened after completion. The flow a user follows:

  1. Open the preview dialog ("Sandbox Preview") and pick a port — the port the agent's server is listening on inside the pod (the field defaults to 3000; the default allowed range is 3000–9000).
  2. Start the preview. The app calls apiClient.startPortForward(runId, port) → POST /api/runs/{runId}/sandbox/port-forward with that port. Despite the historical endpoint name, the Gateway path resolves the bound pod and creates a ClusterIP Service and HTTPRoute.
  3. See it become active. The dialog confirms with "Preview active for port {target_port} on pod {pod_name}" and shows the session id. If the API returned a proxied preview_url, the dialog embeds it in an iframe with an "Open preview" button. A public preview_url is the normal Gateway-path result; the dialog retains an explicit no-URL fallback for deployments without that path.
  4. Stop it when done. Stopping the preview calls apiClient.stopPortForward(runId, sessionId) → DELETE on that session, removing its route and service. You can run more than one preview at a time (up to a per-run cap), each its own port and entry, and stop them individually.

When to use it:

  • You want to look at what the agent built/ran — a running web app, an API the agent stood up, a served artifact — without leaving the run view.
  • You're debugging inside the sandbox and need a live endpoint into the pod for the moment.

What to expect:

  • Kubernetes-only. The Gateway routes into the agent-sandbox controller's pod. On local/dev runs there is no claim pod to forward, so the button doesn't appear — the same "this is a cluster feature" boundary as the pod pill. (Local runs isolate commands with MXC, which has no pod to forward into.)
  • A public capability URL. Anyone with the unguessable HTTPS URL can reach the preview. Do not share it as though it required an Agentweaver login.
  • Scoped to this run's pod. A preview reaches only the run's own sandbox pod, never another run's, and is capped (default 3 per run, 20 globally).
  • Bounded lifetime and retained resources. Project preview lifetime defaults to 1440 minutes (24 hours), used for both initial expiry and the hard cap. Keepalive cannot extend that hard cap. Preview retention can keep the claim/process available after a run ends or suspends; expiry, explicit Stop, or pod loss ends availability. A released/replaced pod may require a fresh preview.

The endpoints and the PortForwardSessionDto fields behind the dialog are in the reference.

Dedicated pages: see the Sandbox browser preview User Guide for the full step-by-step, the Reference for the API, and the Deep Dive for the Gateway control and data paths.

Suspend and resume, from the user's view ​

Pod-per-run uses a hybrid lifecycle: the pod stays warm during active reasoning, and the worker can release it at a workflow suspension when Sandbox:ReleasePodOnSuspend=true. This is a conditional, best-effort release, not a rule that every external wait holds no pod. Assembly and active previews may retain their pod/process and worktree.

What the user experiences across that boundary:

  • At a review/confirmation gate, the run pauses for a decision. Where release is enabled and no retention applies, the worker attempts to release the execution pod while the human decides.
  • When they submit the decision, a released pod is re-claimed and the run is rehydrated from its checkpoint. The worktree is already there (it lives on the shared workspace volume), and the resumed pod gets a fresh run-scoped credential.
  • The pod name may change after resume. Because resume re-claims a (possibly different) warm pod, the pod pill on the agent box can show a different name than before the pause. That is expected and is the visible trace of release-and-rehydrate; the run, its history, and its workspace are continuous.

The coordinator's orchestration loop stays in the worker — only leaf agent turns occupy AgentHost. Its timeline and steering do not depend on retaining a coordinator pod, although assembly preview resources can deliberately stay alive during review.

A debug/low-latency option exists for operators (Sandbox:ReleasePodOnSuspend = false) that keeps the pod warm across a suspension, at the cost of holding capacity. With release enabled, a name change across a pause is normal; with retention, the same name can remain. Pod loss can change placement in either case.

The operator's mental model ​

An operator reasons about pod-per-run as "each run rents a pod for its active bursts, and gives it back when it's waiting." Concretely:

  • Mode is configuration. Sandbox:AgentExecutionMode selects in-process (in-api, the code default) or per-run pods (pod-per-run, selected by the checked-in API/worker manifests). Apply changes through deployment; a mode change does not instantly migrate active turns. Sandbox:ReleasePodOnSuspend (default on) controls ordinary suspension release.
  • Kubernetes owns scheduling; the platform does not pre-gate on quota. More concurrent runs means more claimed AgentHost pods; the agentweaver-agent-host pool keeps 2 run pods pre-warmed. The platform no longer checks quota headroom before a launch — it submits the SandboxClaim and waits for Kubernetes to schedule and bind the pod, so a Pending pod is an expected transient state (a node is freeing up or katapool is autoscaling), not a failure. The namespace ResourceQuota bounds only object counts (pod count, sandbox-claim count, PVCs, storage) — it no longer caps CPU/memory, and those counts are raised deliberately in the manifests, not patched live. See Operations and the reference.
  • The isolation backend is chosen per host. Independently of pod-per-run, every host selects one command-isolation backend at run start and announces it with a sandbox.selected event (backend, isRealIsolation, reason). In-cluster that backend is the Kata-isolated kubernetes-sandbox-claim, whose pods are provisioned by the upstream agent-sandbox controller (installed by scripts/azure/steps/10-create-cluster.mjs); local dev gets processcontainer (MXC, a different local-host runtime) on Windows or linux-bwrap on Linux, falling back to direct (no isolation, shell still runs) only when nothing else is available. The deep dive's executor seam and agent-sandbox controller sections explain how these relate to pod-per-run; backend install/selection is in Sandbox setup.
  • The pod is disposable, not transparently replayable. Durable workspace/checkpoint state supports recovery, but an interrupted A2A turn can fail visibly and require policy-driven redispatch. Retained preview pods can outlive the run; normal cleanup must not be confused with preview expiry.
  • Blast radius is small and visible. Each pod is Kata-isolated, default-deny on egress (model + worker + git only, never the database), and holds no broker key. Inside the pod the boundary goes one level further: model-controlled commands run in a separate executor sidecar container with its own PID namespace and no cluster or cloud identity, so an injected agent cannot see, signal, or read the AgentHost process that holds the run's brokered GitHub token. The pod pill in the UI is the operator's quick "which pod is this run in?" answer.

Diagnostics surface (MCP and runtime) ​

The same facts are available outside the topology view:

  • GET /api/system/runtime answers "are we in Kubernetes, and what is the host pod name?" It is host diagnostics, not a replacement for missing child identity in the coordinator graph.
  • GET /api/runs/{id}/graph carries executionPodName per node — the authoritative "which pod is this run/node executing in?" for a specific run.
  • MCP operations tools expose the same operational health an operator needs around runs (diagnostics, heartbeat, sandbox policy). The MCP surface mirrors the web operations surface fact-for-fact; see the MCP client experience and Operations. Pod naming itself is a presentation detail surfaced primarily in the web topology; the underlying run/pod state is the same the API exposes.

The transport that carries agent turns into the pod is the A2A bridge, which ships on an experimental -preview package line. Operators should treat it as such — pinned and behind the rollback flag — and read the A2A distributed agents experience for what that means in practice.

Edge cases the user may notice ​

  • No pill at all. Local/dev or non-Kubernetes backends never show the pod pill — expected, not a bug.
  • Pill appears slightly after a node starts. The name is shown once the pod is bound and registered; a brief gap before the pill appears is normal.
  • Pill name changes after a pause. Release-on-suspend re-claims a fresh pod on resume, so the name can change across a review gate or a coordinator wait when release occurs.
  • Two runs, two pod names. Different names show distinct recorded pod placements, not proof of every isolation control.
Diagram details and constraints
ElementContract
titleFrom run request to reviewable work
takeawayA run gets isolated execution, visible progress and bounded tools—not an unrestricted host shell.
group-title0ADMIT AND PREPARE
group-title1EXECUTE AND RETURN EVIDENCE
Authorized runAuthorized run
Authorized runUser requests repository work
Authorized runProvider acceptance first
Authorized runAPI owns run lifecycle
Authorized runProject role required
SandboxClaimSandboxClaim
SandboxClaimBind a warm AgentHost pod
SandboxClaimResolve claim-bound pod
SandboxClaimConfigure identity once
SandboxClaimClaim state is shared
AgentHostAgentHost
AgentHostModel and governed tools
AgentHostA2A authenticated turns
AgentHostRun-scoped workspace
AgentHostCopilot OR BYOK
Run experienceRun experience
Run experienceEvents and human decisions
Run experienceProgress / tool evidence
Run experienceApprove consequential work
Run experienceApproval ≠ policy bypass
Isolated workspaceIsolated workspace
Isolated workspaceFile and shell results
Isolated workspaceContain paths and processes
Isolated workspaceBound / redact tool output
Isolated workspaceNetwork policy also applies
Preview or reviewPreview or review
Preview or reviewInspect resulting work
Preview or reviewPreview needs publication proof
Preview or reviewReview artifacts before merge
Preview or reviewRelease / reap compute
relation-01 start
relation-12 configure
relation-23 tools
relation-34 events / approvals
relation-45 work artifacts
assuranceCredentials arrive via /configure, not ambient stores. Approval, policy and isolation remain independent checks.
assurance-0-labelWorkspace ownership
assurance-0-factChildren use isolated worktrees.
assurance-0-sourceRunOrchestrator.cs
assurance-1-labelPod observation
assurance-1-factPod telemetry is not branch ownership.
assurance-1-sourcesandbox-pod-execution.md
assurance-2-labelPublication evidence
assurance-2-factPreview readiness proves public HTTPS.
assurance-2-sourceSandboxPreviewPublicationTests.cs
n0Provider acceptance first; API owns run lifecycle
n1Resolve claim-bound pod; Configure identity once
n2A2A authenticated turns; Run-scoped workspace
n3Progress / tool evidence; Approve consequential work
n4Contain paths and processes; Bound / redact tool output
n5Preview needs publication proof; Review artifacts before merge
groupsADMIT AND PREPARE; EXECUTE AND RETURN EVIDENCE
Diagram details and constraints
ElementContract
titleA durable run can outlive its pod
takeawayA review wait may release compute; preview and assembly retention are explicit exceptions.
group-title-0ACTIVE WORK AND WAIT
group-title-1RESUME OR RETAIN
Active leafActive leaf
Active leafAgentHost returns output
Active leafpod -> worker -> timeline
Active leafWorker records events and owns workflow progression.
Worker gateWorker gate
Worker gateHuman decision requested
Worker gatecheckpoint-backed wait
Worker gatePersist resumable state; pause watchdog accounting.
Release decisionRelease decision
Release decisionPolicy and lifecycle checks
Release decisionpod-per-run + enabled
Release decisionRelease requires support and no active retention exception.
Resumed workResumed work
Resumed workContinue from saved state
Resumed workevents through worker
Resumed workA released pod may be replaced; its name can change.
Durable checkpointDurable checkpoint
Durable checkpointSession and workspace
Durable checkpointauthorized decision
Durable checkpointLoad resumable state; claim/configure if released.
Release or retainRelease or retain
Release or retainClaim deletion / keep pod
Release or retainpreview / assembly exception
Release or retainActive preview defers deletion; release failures are logged.
e0reach gate
e1evaluate
e2apply
e3checkpoint
e4resume
noteHuman waiting does not imply zero retained compute. This is not arbitrary failed-turn rehydration.
n0Worker records events and owns workflow progression.
n1Persist resumable state; pause watchdog accounting.
n2Release requires support and no active retention exception.
n3A released pod may be replaced; its name can change.
n4Load resumable state; claim/configure if released.
n5Active preview defers deletion; release failures are logged.
groupsACTIVE WORK AND WAIT; RESUME OR RETAIN
Diagram details and constraints
ElementContract
titlePreview readiness follows the public path
takeawayProvision the route, then probe its exact HTTPS URL; object creation alone is not ready.
group-title0CONTROL: PROVISION + PROBE
group-title1GATEWAY DATA PATH
Preview APIPreview API
Preview APIResolve bound SandboxClaim
Preview APIPatch run selector on pod
Preview APICreate Service + HTTPRoute
Preview APIState from cluster, not cache
Publication probePublication probe
Publication probeExact generated HTTPS URL
Publication probeWait for managed DNS
Publication probeCheck Gateway + application
Publication probeOnly then return ready
Browser previewBrowser preview
Browser previewOpen the returned URL
Browser previewRun-scoped capability host
Browser previewKeepalive via API
Browser previewIframe: no-referrer
Preview GatewayPreview Gateway
Preview GatewaySeparate shared Gateway
Preview GatewayHTTPS host match
Preview GatewayHTTPRoute selects Service
Preview GatewayNot API port-forward
ClusterIP ServiceClusterIP Service
ClusterIP ServicePer-preview target selector
ClusterIP ServiceService :80 → public port
ClusterIP ServiceRoutes to bound sandbox pod
ClusterIP ServiceAllowed ports 3000–9000
Sandbox preview appSandbox preview app
Sandbox preview appAgentHost pod-local path
Sandbox preview appLive preview: TCP forwarder
Sandbox preview app0.0.0.0 → loopback app
Sandbox preview appManual: chosen target port
relation-01 after create
relation-12 ready URL
relation-23 HTTPS probe
relation-34 HTTPS
relation-45 route
relation-56 public port
assuranceNo API → pod TCP readiness probe. Publication failure rolls back; DNS convergence has a bounded retry window.
assurance-0-labelPublic readiness
assurance-0-factProbe the exact generated HTTPS URL.
assurance-0-sourceSandboxPreviewService.cs
assurance-1-labelRollback on failure
assurance-1-factUnpublish failed preview resources.
assurance-1-sourceSandboxPreviewPublicationTests.cs
assurance-2-labelSeparate ingress
assurance-2-factDNS managed externally, not by API.
assurance-2-sourcegateway-preview.yaml
n0Patch run selector on pod; Create Service + HTTPRoute
n1Wait for managed DNS; Check Gateway + application
n2Run-scoped capability host; Keepalive via API
n3HTTPS host match; HTTPRoute selects Service
n4Service :80 → public port; Routes to bound sandbox pod
n5Live preview: TCP forwarder; 0.0.0.0 → loopback app
groupsCONTROL: PROVISION + PROBE; GATEWAY DATA PATH