Sandbox pod execution experience
This page is about what a person sees and feels when Agentweaver runs each agent in its own per-run sandbox pod — both the user watching a run in the web UI and the operator reasoning about the cluster. For the design logic see the Sandbox pod execution deep dive; for flags, identity, and the token mechanism see the Sandbox pods reference.
Related journeys: Runs, board & watch, Coordinator & orchestration, Operations, and the A2A distributed agents experience.
Mental model
The most important thing to feel about pod-per-run is that almost nothing changes in how a run looks. The board, live timeline, coordinator topology, and review gates retain their existing roles. What changes is where the work physically runs: each run's agent now executes inside its own Kata-isolated pod instead of inside the worker process. A pod name on the agent box shows recorded placement. Provisioning delays and remote-turn failures can also be visible.
A well-behaved run should make the user barely notice the sandbox. The pod pill is a placement signal, not independent proof of the sandbox's full isolation or credential posture.
What the pod pill is
When the backend runs inside Kubernetes, each agent box in a run's topology shows a small pod pill — a compact, monospace pill with a server icon and the executing pod's name (for example agentweaver-api-abc123 or agent-pod-worker-7). Hovering it shows the tooltip "Executing in pod {name}", and it is exposed to assistive technology with the same label.
What the user can take from it:
- Where this agent is actually running. Under pod-per-run, the pill on a node is that run's sandbox pod, so two concurrently-running agents can show two different pod names — a direct, visible cue that their work is isolated from each other.
- Which placement was recorded. Different pod names distinguish execution locations; configuration and cluster policy, not the pill alone, establish isolation.
The pill is Kubernetes-only and quiet by design:
- on local/dev runs, or anywhere the backend is not in Kubernetes, no pill is shown — the UI stays clean and there is nothing to explain;
- if the pod name is not yet known (the claim has not bound), no pill is shown rather than a placeholder; and
- if the runtime probe fails, the UI degrades silently to no pill rather than erroring.
Where the value comes from: the coordinator graph passes each node's executionPodName directly to the pod indicator. A missing child-pod identity stays absent; it is not replaced with the API replica's name. The binding is persisted in the shared run-event log so another replica can serve the recorded placement. The data behind it is in the reference.
What happens during a run
From the user's side, submitting and watching a run feels identical to before:
- They submit work (or a coordinator goal) and watch the board / topology as usual.
- Agent boxes light up as running, stream tokens and status live, and — on Kubernetes — show their pod pill.
- The worker records pod-produced events and serves the existing timeline. Provisioning delays or an interrupted remote turn can still become visible.
- When the run completes, the box settles into its terminal state as usual.
Under the hood the heavy work — the model session, the tools, the shell — is happening in the pod, and the worker is relaying its stream into the same timeline. The user does not have to know that; the only new thing on screen is the pod name.
Sandbox preview: reaching a server inside the pod
Sometimes an agent starts a server inside its sandbox pod — a dev server, a freshly built app, a debug endpoint — and a person wants to actually open it and look. Because the pod is isolated, that server is not reachable by default. The sandbox browser preview publishes a run-scoped HTTPS capability URL through the preview Gateway; the API provisions the route, rather than proxying app traffic.
A Preview Sandbox button appears on the coordinator run page only when the run is using the Kubernetes sandbox (sandboxBackend === 'kubernetes-sandbox-claim', read from the run's sandbox.selected event) and a preview lifecycle state or existing preview session is present. The current button is not gated simply on whether the run is active; retained previews can be opened after completion. The flow a user follows:
- Open the preview dialog ("Sandbox Preview") and pick a port — the port the agent's server is listening on inside the pod (the field defaults to
3000; the default allowed range is3000–9000). - Start the preview. The app calls
apiClient.startPortForward(runId, port)→POST /api/runs/{runId}/sandbox/port-forwardwith that port. Despite the historical endpoint name, the Gateway path resolves the bound pod and creates a ClusterIP Service and HTTPRoute. - See it become active. The dialog confirms with "Preview active for port {target_port} on pod {pod_name}" and shows the session id. If the API returned a proxied
preview_url, the dialog embeds it in an iframe with an "Open preview" button. A publicpreview_urlis the normal Gateway-path result; the dialog retains an explicit no-URL fallback for deployments without that path. - Stop it when done. Stopping the preview calls
apiClient.stopPortForward(runId, sessionId)→DELETEon that session, removing its route and service. You can run more than one preview at a time (up to a per-run cap), each its own port and entry, and stop them individually.
When to use it:
- You want to look at what the agent built/ran — a running web app, an API the agent stood up, a served artifact — without leaving the run view.
- You're debugging inside the sandbox and need a live endpoint into the pod for the moment.
What to expect:
- Kubernetes-only. The Gateway routes into the agent-sandbox controller's pod. On local/dev runs there is no claim pod to forward, so the button doesn't appear — the same "this is a cluster feature" boundary as the pod pill. (Local runs isolate commands with MXC, which has no pod to forward into.)
- A public capability URL. Anyone with the unguessable HTTPS URL can reach the preview. Do not share it as though it required an Agentweaver login.
- Scoped to this run's pod. A preview reaches only the run's own sandbox pod, never another run's, and is capped (default 3 per run, 20 globally).
- Bounded lifetime and retained resources. Project preview lifetime defaults to 1440 minutes (24 hours), used for both initial expiry and the hard cap. Keepalive cannot extend that hard cap. Preview retention can keep the claim/process available after a run ends or suspends; expiry, explicit Stop, or pod loss ends availability. A released/replaced pod may require a fresh preview.
The endpoints and the PortForwardSessionDto fields behind the dialog are in the reference.
Dedicated pages: see the Sandbox browser preview User Guide for the full step-by-step, the Reference for the API, and the Deep Dive for the Gateway control and data paths.
Suspend and resume, from the user's view
Pod-per-run uses a hybrid lifecycle: the pod stays warm during active reasoning, and the worker can release it at a workflow suspension when Sandbox:ReleasePodOnSuspend=true. This is a conditional, best-effort release, not a rule that every external wait holds no pod. Assembly and active previews may retain their pod/process and worktree.
What the user experiences across that boundary:
- At a review/confirmation gate, the run pauses for a decision. Where release is enabled and no retention applies, the worker attempts to release the execution pod while the human decides.
- When they submit the decision, a released pod is re-claimed and the run is rehydrated from its checkpoint. The worktree is already there (it lives on the shared workspace volume), and the resumed pod gets a fresh run-scoped credential.
- The pod name may change after resume. Because resume re-claims a (possibly different) warm pod, the pod pill on the agent box can show a different name than before the pause. That is expected and is the visible trace of release-and-rehydrate; the run, its history, and its workspace are continuous.
The coordinator's orchestration loop stays in the worker — only leaf agent turns occupy AgentHost. Its timeline and steering do not depend on retaining a coordinator pod, although assembly preview resources can deliberately stay alive during review.
A debug/low-latency option exists for operators (Sandbox:ReleasePodOnSuspend = false) that keeps the pod warm across a suspension, at the cost of holding capacity. With release enabled, a name change across a pause is normal; with retention, the same name can remain. Pod loss can change placement in either case.
The operator's mental model
An operator reasons about pod-per-run as "each run rents a pod for its active bursts, and gives it back when it's waiting." Concretely:
- Mode is configuration.
Sandbox:AgentExecutionModeselects in-process (in-api, the code default) or per-run pods (pod-per-run, selected by the checked-in API/worker manifests). Apply changes through deployment; a mode change does not instantly migrate active turns.Sandbox:ReleasePodOnSuspend(default on) controls ordinary suspension release. - Kubernetes owns scheduling; the platform does not pre-gate on quota. More concurrent runs means more claimed AgentHost pods; the
agentweaver-agent-hostpool keeps 2 run pods pre-warmed. The platform no longer checks quota headroom before a launch — it submits theSandboxClaimand waits for Kubernetes to schedule and bind the pod, so a Pending pod is an expected transient state (a node is freeing up orkatapoolis autoscaling), not a failure. The namespaceResourceQuotabounds only object counts (pod count, sandbox-claim count, PVCs, storage) — it no longer caps CPU/memory, and those counts are raised deliberately in the manifests, not patched live. See Operations and the reference. - The isolation backend is chosen per host. Independently of pod-per-run, every host selects one command-isolation backend at run start and announces it with a
sandbox.selectedevent (backend,isRealIsolation,reason). In-cluster that backend is the Kata-isolatedkubernetes-sandbox-claim, whose pods are provisioned by the upstream agent-sandbox controller (installed byscripts/azure/steps/10-create-cluster.mjs); local dev getsprocesscontainer(MXC, a different local-host runtime) on Windows orlinux-bwrapon Linux, falling back todirect(no isolation, shell still runs) only when nothing else is available. The deep dive's executor seam and agent-sandbox controller sections explain how these relate to pod-per-run; backend install/selection is in Sandbox setup. - The pod is disposable, not transparently replayable. Durable workspace/checkpoint state supports recovery, but an interrupted A2A turn can fail visibly and require policy-driven redispatch. Retained preview pods can outlive the run; normal cleanup must not be confused with preview expiry.
- Blast radius is small and visible. Each pod is Kata-isolated, default-deny on egress (model + worker + git only, never the database), and holds no broker key. Inside the pod the boundary goes one level further: model-controlled commands run in a separate executor sidecar container with its own PID namespace and no cluster or cloud identity, so an injected agent cannot see, signal, or read the AgentHost process that holds the run's brokered GitHub token. The pod pill in the UI is the operator's quick "which pod is this run in?" answer.
Diagnostics surface (MCP and runtime)
The same facts are available outside the topology view:
GET /api/system/runtimeanswers "are we in Kubernetes, and what is the host pod name?" It is host diagnostics, not a replacement for missing child identity in the coordinator graph.GET /api/runs/{id}/graphcarriesexecutionPodNameper node — the authoritative "which pod is this run/node executing in?" for a specific run.- MCP operations tools expose the same operational health an operator needs around runs (diagnostics, heartbeat, sandbox policy). The MCP surface mirrors the web operations surface fact-for-fact; see the MCP client experience and Operations. Pod naming itself is a presentation detail surfaced primarily in the web topology; the underlying run/pod state is the same the API exposes.
The transport that carries agent turns into the pod is the A2A bridge, which ships on an experimental -preview package line. Operators should treat it as such — pinned and behind the rollback flag — and read the A2A distributed agents experience for what that means in practice.
Edge cases the user may notice
- No pill at all. Local/dev or non-Kubernetes backends never show the pod pill — expected, not a bug.
- Pill appears slightly after a node starts. The name is shown once the pod is bound and registered; a brief gap before the pill appears is normal.
- Pill name changes after a pause. Release-on-suspend re-claims a fresh pod on resume, so the name can change across a review gate or a coordinator wait when release occurs.
- Two runs, two pod names. Different names show distinct recorded pod placements, not proof of every isolation control.
Related reading
- Sandbox pod execution deep dive — the why and the logic.
- Sandbox pods reference — flags, identity/quota, token injection, naming.
- Runs, board & watch and Coordinator & orchestration — the journeys this annotates.
- Operations — health, heartbeat, and sandbox policy surfaces.
- A2A distributed agents experience — the
-previewtransport behind it.
Diagram details and constraints
| Element | Contract |
|---|---|
| title | From run request to reviewable work |
| takeaway | A run gets isolated execution, visible progress and bounded tools—not an unrestricted host shell. |
| group-title0 | ADMIT AND PREPARE |
| group-title1 | EXECUTE AND RETURN EVIDENCE |
| Authorized run | Authorized run |
| Authorized run | User requests repository work |
| Authorized run | Provider acceptance first |
| Authorized run | API owns run lifecycle |
| Authorized run | Project role required |
| SandboxClaim | SandboxClaim |
| SandboxClaim | Bind a warm AgentHost pod |
| SandboxClaim | Resolve claim-bound pod |
| SandboxClaim | Configure identity once |
| SandboxClaim | Claim state is shared |
| AgentHost | AgentHost |
| AgentHost | Model and governed tools |
| AgentHost | A2A authenticated turns |
| AgentHost | Run-scoped workspace |
| AgentHost | Copilot OR BYOK |
| Run experience | Run experience |
| Run experience | Events and human decisions |
| Run experience | Progress / tool evidence |
| Run experience | Approve consequential work |
| Run experience | Approval ≠ policy bypass |
| Isolated workspace | Isolated workspace |
| Isolated workspace | File and shell results |
| Isolated workspace | Contain paths and processes |
| Isolated workspace | Bound / redact tool output |
| Isolated workspace | Network policy also applies |
| Preview or review | Preview or review |
| Preview or review | Inspect resulting work |
| Preview or review | Preview needs publication proof |
| Preview or review | Review artifacts before merge |
| Preview or review | Release / reap compute |
| relation-0 | 1 start |
| relation-1 | 2 configure |
| relation-2 | 3 tools |
| relation-3 | 4 events / approvals |
| relation-4 | 5 work artifacts |
| assurance | Credentials arrive via /configure, not ambient stores. Approval, policy and isolation remain independent checks. |
| assurance-0-label | Workspace ownership |
| assurance-0-fact | Children use isolated worktrees. |
| assurance-0-source | RunOrchestrator.cs |
| assurance-1-label | Pod observation |
| assurance-1-fact | Pod telemetry is not branch ownership. |
| assurance-1-source | sandbox-pod-execution.md |
| assurance-2-label | Publication evidence |
| assurance-2-fact | Preview readiness proves public HTTPS. |
| assurance-2-source | SandboxPreviewPublicationTests.cs |
| n0 | Provider acceptance first; API owns run lifecycle |
| n1 | Resolve claim-bound pod; Configure identity once |
| n2 | A2A authenticated turns; Run-scoped workspace |
| n3 | Progress / tool evidence; Approve consequential work |
| n4 | Contain paths and processes; Bound / redact tool output |
| n5 | Preview needs publication proof; Review artifacts before merge |
| groups | ADMIT AND PREPARE; EXECUTE AND RETURN EVIDENCE |
Diagram details and constraints
| Element | Contract |
|---|---|
| title | A durable run can outlive its pod |
| takeaway | A review wait may release compute; preview and assembly retention are explicit exceptions. |
| group-title-0 | ACTIVE WORK AND WAIT |
| group-title-1 | RESUME OR RETAIN |
| Active leaf | Active leaf |
| Active leaf | AgentHost returns output |
| Active leaf | pod -> worker -> timeline |
| Active leaf | Worker records events and owns workflow progression. |
| Worker gate | Worker gate |
| Worker gate | Human decision requested |
| Worker gate | checkpoint-backed wait |
| Worker gate | Persist resumable state; pause watchdog accounting. |
| Release decision | Release decision |
| Release decision | Policy and lifecycle checks |
| Release decision | pod-per-run + enabled |
| Release decision | Release requires support and no active retention exception. |
| Resumed work | Resumed work |
| Resumed work | Continue from saved state |
| Resumed work | events through worker |
| Resumed work | A released pod may be replaced; its name can change. |
| Durable checkpoint | Durable checkpoint |
| Durable checkpoint | Session and workspace |
| Durable checkpoint | authorized decision |
| Durable checkpoint | Load resumable state; claim/configure if released. |
| Release or retain | Release or retain |
| Release or retain | Claim deletion / keep pod |
| Release or retain | preview / assembly exception |
| Release or retain | Active preview defers deletion; release failures are logged. |
| e0 | reach gate |
| e1 | evaluate |
| e2 | apply |
| e3 | checkpoint |
| e4 | resume |
| note | Human waiting does not imply zero retained compute. This is not arbitrary failed-turn rehydration. |
| n0 | Worker records events and owns workflow progression. |
| n1 | Persist resumable state; pause watchdog accounting. |
| n2 | Release requires support and no active retention exception. |
| n3 | A released pod may be replaced; its name can change. |
| n4 | Load resumable state; claim/configure if released. |
| n5 | Active preview defers deletion; release failures are logged. |
| groups | ACTIVE WORK AND WAIT; RESUME OR RETAIN |
Diagram details and constraints
| Element | Contract |
|---|---|
| title | Preview readiness follows the public path |
| takeaway | Provision the route, then probe its exact HTTPS URL; object creation alone is not ready. |
| group-title0 | CONTROL: PROVISION + PROBE |
| group-title1 | GATEWAY DATA PATH |
| Preview API | Preview API |
| Preview API | Resolve bound SandboxClaim |
| Preview API | Patch run selector on pod |
| Preview API | Create Service + HTTPRoute |
| Preview API | State from cluster, not cache |
| Publication probe | Publication probe |
| Publication probe | Exact generated HTTPS URL |
| Publication probe | Wait for managed DNS |
| Publication probe | Check Gateway + application |
| Publication probe | Only then return ready |
| Browser preview | Browser preview |
| Browser preview | Open the returned URL |
| Browser preview | Run-scoped capability host |
| Browser preview | Keepalive via API |
| Browser preview | Iframe: no-referrer |
| Preview Gateway | Preview Gateway |
| Preview Gateway | Separate shared Gateway |
| Preview Gateway | HTTPS host match |
| Preview Gateway | HTTPRoute selects Service |
| Preview Gateway | Not API port-forward |
| ClusterIP Service | ClusterIP Service |
| ClusterIP Service | Per-preview target selector |
| ClusterIP Service | Service :80 → public port |
| ClusterIP Service | Routes to bound sandbox pod |
| ClusterIP Service | Allowed ports 3000–9000 |
| Sandbox preview app | Sandbox preview app |
| Sandbox preview app | AgentHost pod-local path |
| Sandbox preview app | Live preview: TCP forwarder |
| Sandbox preview app | 0.0.0.0 → loopback app |
| Sandbox preview app | Manual: chosen target port |
| relation-0 | 1 after create |
| relation-1 | 2 ready URL |
| relation-2 | 3 HTTPS probe |
| relation-3 | 4 HTTPS |
| relation-4 | 5 route |
| relation-5 | 6 public port |
| assurance | No API → pod TCP readiness probe. Publication failure rolls back; DNS convergence has a bounded retry window. |
| assurance-0-label | Public readiness |
| assurance-0-fact | Probe the exact generated HTTPS URL. |
| assurance-0-source | SandboxPreviewService.cs |
| assurance-1-label | Rollback on failure |
| assurance-1-fact | Unpublish failed preview resources. |
| assurance-1-source | SandboxPreviewPublicationTests.cs |
| assurance-2-label | Separate ingress |
| assurance-2-fact | DNS managed externally, not by API. |
| assurance-2-source | gateway-preview.yaml |
| n0 | Patch run selector on pod; Create Service + HTTPRoute |
| n1 | Wait for managed DNS; Check Gateway + application |
| n2 | Run-scoped capability host; Keepalive via API |
| n3 | HTTPS host match; HTTPRoute selects Service |
| n4 | Service :80 → public port; Routes to bound sandbox pod |
| n5 | Live preview: TCP forwarder; 0.0.0.0 → loopback app |
| groups | CONTROL: PROVISION + PROBE; GATEWAY DATA PATH |
