Sandbox pods reference
Exhaustive reference for pod-per-run sandbox execution: configuration flags, pod identity and quota, run-scoped GitHub token injection, pod naming, and the security properties of the model. For the reasoning behind these mechanics, see the Sandbox pod execution deep dive; for the operator/user view, see the Sandbox pod execution experience.
This page documents the sandbox-pod execution surface (where the agent turn runs). The broader sandbox isolation model — filesystem containment, governance, executor selection, and claim lifecycle — is the Sandbox deep dive, and operator install/config is Sandbox setup.
Configuration flags
| Flag | Values | Default | Effect |
|---|---|---|---|
Sandbox:AgentExecutionMode | in-api, pod-per-run | in-api | in-api runs the agent turn in-process in the API/worker (today's behavior, the rollback path). pod-per-run relocates each run's agent turn into its own Kata-isolated sandbox pod via the A2A bridge. |
Sandbox:ReleasePodOnSuspend | true, false | true | When pod-per-run is active and the workflow graph suspends on an external gate (a HITL/review RequestPort, or the coordinator idling while it awaits child runs), true checkpoints the run and releases the claim and deletes the used pod so the pool replenishes capacity; an active preview can defer release. false keeps the pod warm across the suspension for low-latency resume or debugging, at the cost of held capacity. |
Sandbox:Kubernetes:AgentHostClaimCreationGraceSeconds | Positive integer seconds | 300 | Minimum age before the orphan reaper may delete an AgentHost claim that is absent from the active-run map. The effective grace is the larger of this value and Sandbox:Kubernetes:AgentHostReadyTimeoutSeconds + 30 seconds. |
Sandbox:Kubernetes:AgentHostProvisioningTimeoutSeconds | Positive integer seconds | 600 | Maximum time for an AgentHost SandboxClaim to bind. A scheduling wait beyond this limit fails the launch and releases the claim; the failure includes the pod's PodScheduled=False diagnosis when available. This is separate from the AgentHost readiness timeout after binding. |
AgentHost:ExecutionScratchRoot | Absolute path | /local-workspace | Root of the disk-backed emptyDir used for pod-local execution workspaces and package caches. |
AgentHost:ExecutionScratchMinimumFreeBytes | Non-negative integer bytes | 8589934592 (8 GiB) | Minimum available scratch space required before AgentHost prepares a local workspace. Failure returns typed reason insufficient_ephemeral_storage. |
Coordinator:AssemblyBuildTestTimeoutMinutes | Positive number | 20 | Total assembly Build/Test wall-clock limit. Expiry cancels the gate and releases its retained AgentHost claim. |
Coordinator:AssemblyBuildTestStallTimeoutMinutes | Positive number | 12 | Maximum interval without a forwarded Build/Test run event before the stall watchdog fails the gate. |
Flag semantics
pod-per-runis the only value that activates the bridge. Any other value (thein-apidefault) keeps execution in-process. There is no separate "pod-per-turn" mode — granularity withinpod-per-runis the hybrid model (warm across consecutive turns, release on suspend), governed bySandbox:ReleasePodOnSuspend, not by a distinct execution-mode value.ReleasePodOnSuspendonly matters underpod-per-run. It is a tuning sub-flag; it never changes the execution-mode value. The release is internal behavior ofpod-per-run.- Rollback is a flag flip, not a redeploy. Setting
Sandbox:AgentExecutionMode=in-apirestores in-process execution immediately. This is the documented mitigation for any instability in the-previewA2A transport — there is no alternate wire transport to deploy. See the A2A reference for the transport's preview status and pinning.
AgentHost receives only a live, immutable capability credential redeemed for the configured run and purpose through
/configure. It has no Key Vault, CSI, shared-filesystem, or ambient-token fallback.
Pod identity and quota
A pod-per-run sandbox is the same Kata-isolated pod shape the sandbox subsystem already uses, claimed from a warm pool, but now hosting the full agent (worker agents and the coordinator's own agent turns) rather than only ad-hoc shell commands.
| Property | Value / behavior |
|---|---|
| Runtime class | kata-vm-isolation — a VM boundary around the container, so each run's secret and execution live inside a per-run microVM and are destroyed with it. |
| Identity | Dedicated sandbox service account federated to agentweaver-agenthost-identity, a managed identity with no Key Vault role assignments (issue #471). Workload identity (federated OIDC) projects only the narrowly-scoped workload-identity token volume — not the full Kubernetes API service-account token — but it grants no vault access, so the sandbox cannot read any user's secrets. |
| Cluster API access | Current pod infrastructure enables service-account automount; the model-execution sidecar masks /var/run/secrets/kubernetes.io/serviceaccount. The whole pod is not universally tokenless. |
| Provisioning | Claimed from a warm pool via a SandboxClaim; the executor waits until the claim is bound to a concrete pod. AgentHost uses the shared agentweaver-agent-host pool (replicas: 2), then receives per-run context through POST /configure before /healthz is expected to become ready. No separate per-run template or per-run warm pool is created for AgentHost. A claim that stays unbound (pod Pending) while Kubernetes schedules is a legitimate transient wait — there is no app-side capacity pre-check — surfaced on the child run's stream via sandbox.provisioning_pending heartbeats including the scheduler reason when available. It fails and releases its claim after the provisioning limit. |
| AgentHost readiness gate | Warm AgentHost pods start in standby. After binding, the executor calls POST /configure with run identity, a live Copilot capability or BYOK configuration, separate purpose-scoped credentials, and the execution-workspace descriptor, then polls GET {scheme}://{podIP}:8088/healthz (bounded Sandbox:Kubernetes:AgentHostReadyTimeoutSeconds, default 90s; …ReadyPollIntervalMs, default 1000) before the first A2A turn. /configure is excluded from readiness and returns 409 if called again. The a2a-sandbox-pod HttpClient additionally retries connection-refused only. |
| Transient API resilience | The idempotent claim create and the bind/IP polls (WaitForBoundAsync, GetPodIpAsync) retry transient Kubernetes API faults up to MaxK8sAttempts (3 total) with exponential backoff + jitter (ExecuteK8sWithRetryAsync): connection resets (SocketException 104/IOException/HttpRequestException), 429/5xx, and HttpClient timeouts. 409 Conflict is not treated as transient — it is attempt-aware to preserve idempotency (a retry-409 = our own create that committed before a reset, so the claim is configured, not reused). Caller cancellation is never retried. The non-idempotent POST /configure is intentionally excluded (issue #230). |
| A2A turn authentication | Run launch generates a 256-bit random turn bearer token, sends it to the claimed warm pod in POST /configure, and registers it in IAgentHostTurnTokenRegistry. RemoteAgentProxy sends Authorization: Bearer {token} on message:stream; each pod accepts only its configured run token. |
| Tool-approval return path | When the API-side durable approval gate reports Unknown, pod-per-run mode forwards the grant/deny to the owning AgentHost pod's authenticated root endpoint so its in-memory gate can resolve. |
| Resources | AgentHost: requests 300m CPU/1Gi, limits 800m/2Gi. Execution sidecar: requests 700m/2Gi, limits 1200m/4Gi. Each requests 1Gi and limits 4Gi ephemeral storage; shared execution scratch is capped at 8Gi. |
| Quota | Namespace ResourceQuota (k8s/base/quota.yaml) bounds only object counts — pod count, sandbox-claim count, PVCs, and storage. It no longer caps CPU/memory: Kubernetes schedules on pod requests and the cluster autoscaler owns headroom, so a Pending pod waits for the pool to scale rather than being rejected on admission (issue #217). The object-count caps are raised deliberately via a reviewed manifest change, never a live patch. |
| Lifetime | Bounded by the run and the claim TTL. Under the hybrid model, a pod is released on suspend and a fresh pod is re-claimed on resume; a pod can outlive execution while an active preview retains it; release and orphan cleanup resume after durable retention evidence expires. |
| Egress | Default-deny with explicit API/MCP/DNS paths and public HTTPS excluding private/link-local ranges; not a per-run Git-host-only allowlist. No direct PostgreSQL access. |
| Storage | Mounts the shared workspace volume plus a dedicated disk-backed execution-scratch emptyDir at /local-workspace (sizeLimit: 8Gi) for pod-local execution. Assembly Build/Test and preview use LocalReadOnly; implementation turns use LocalWritable and publish through the verified Git write-back flow. Existing disk-backed tmp and home emptyDirs remain separate. |
Orphan reaper creation grace
An AgentHost claim missing from the active-run map is not reaped while its Kubernetes creationTimestamp is inside the effective creation-grace window. This keeps a newly bound claim alive through the readiness wait (AgentHostReadyTimeoutSeconds, default 90 seconds); a missing or unparseable timestamp receives no grace and remains eligible for cleanup.
Run-scoped GitHub capability delivery
A pod-per-run sandbox receives only capability credentials tied to the run's immutable snapshots. RunGitHubCapabilitySnapshotLifecycle captures snapshots before launch and gives retries/resumes fresh references to the inherited capability. The API's GitHubCapabilityBroker fences the selected UnattendedCopilot or UnattendedRepository snapshot before and after redemption, then bounds the credential expiry.
In GitHub Copilot mode, KubernetesSandboxExecutor requires a live Copilot credential for the exact run. It transfers that credential in-memory through the one-time /configure call. In BYOK mode, the sandbox resolves its configured provider separately and does not require or use copilotCredential. AgentHostGitHubCapabilityCredentialProvider rejects credentials for a different run or past expiry. AgentHost does not read Key Vault, CSI mounts, shared filesystem tokens, user token stores, or configuration credentials.
/configure field | Required | Meaning |
|---|---|---|
runId | Yes | Configured run identity. |
copilotCredential | GitHub Copilot mode only | Opaque snapshot reference, credential, and bounded expiry for that run's unattended Copilot capability. BYOK mode does not use this field. |
repositoryAccessToken | No | Separately redeemed repository capability for narrowly-scoped Git/GitHub operations. |
turnBearerToken | No | A2A turn authorization token, distinct from the GitHub capability. |
Credentials are never logged or persisted. Missing, revoked, expired, or purpose-mismatched snapshots fail closed before the pod becomes ready.
A2A turn bearer token
The A2A turn endpoint has a separate per-run bearer token from the GitHub user token above:
KubernetesSandboxExecutorcreates 32 random bytes (256bits) at AgentHost run launch.- The token is sent to the claimed warm pod in
POST /configureand stored inAgentHostRuntimeState. - The same token is stored in
IAgentHostTurnTokenRegistryfor the owning run. RemoteAgentProxyreads the registry and sendsAuthorization: Bearer {token}on all calls toPOST /a2a/agent/v1/message:stream.AgentHostrejects turn requests whose header does not exactly match its ownAgentHostOptions.TurnBearerToken.
This is application-layer auth on top of the A2A NetworkPolicy/mTLS boundary. The important blast-radius property is that a stolen token from one run cannot be reused against another run's pod.
Tool-approval forwarding endpoints
These are internal API-to-AgentHost routes, not public client endpoints. The public caller continues to use /api/runs/{id}/tool-approvals and /api/runs/{id}/tool-denials.
| Method | AgentHost path | Body | Purpose |
|---|---|---|---|
POST | /tool-approvals | runId, requestId, scope | Grant the pod-local pending request. Unknown scope values use once; always is pod/run-scoped and does not survive restart. |
POST | /tool-denials | runId, requestId | Deny the pod-local pending request. |
Both routes accept the same pod-root bearer authorization used by PreviewRunner controls: either the configured turn bearer or the per-run previewRunnerCredential. A mismatched runId returns 409 state: "run_mismatch".
| AgentHost response | Meaning |
|---|---|
200 with resolved: true | State is approved, denied, or expired |
404 with state: "unknown" | The pod-local gate does not know the request |
409 with state: "pending" | The request remains pending |
401 | The bearer did not match the configured pod credentials |
The API locates the pod with IAgentHostOriginResolver, calls it through the a2a-sandbox-pod client, and caps the decision call at 10 seconds. Missing origins, timeouts, transport failures, 5xx responses, and invalid responses surface publicly as 503 state: "agenthost_unreachable". Terminal forwards cause the API to emit tool.approval_resolved for the owning run.
The credential's secret-store key is derived by PreviewRunnerCredential.SecretKey(runId) with the prefix preview-runner-cred--; KubernetesSandboxExecutor mints it, persists it, and delivers its value in-memory through /configure. Key Vault cleanup uses soft delete rather than purge. If the same run launches again while that deterministic key is deleted but recoverable, the API recovers the key, waits up to 30 seconds for it to become writable, and replaces the recovered value with a fresh credential. Concurrent recovery attempts converge on the same active key. Secret values are never included in recovery errors or logs, and terminal cleanup continues to bound the credential lifetime to the backing pod without weakening Key Vault purge protection.
Sources: apps/Agentweaver.AgentHost/Program.cs:287-288,486-588, apps/Agentweaver.Api/Sandbox/AgentHostApprovalHttpClient.cs:28-112, apps/Agentweaver.Api/Endpoints/RunEndpoints.cs:2590-2718, apps/Agentweaver.Api/Sandbox/Preview/PreviewRunnerCredential.cs:22-35, and apps/Agentweaver.Api/Sandbox/KubernetesSandboxExecutor.cs:706-759, apps/Agentweaver.Api/Auth/KeyVaultSecretStore.cs, and apps/Agentweaver.Api/Auth/KeyVaultRecoverableSecretWriter.cs.
Pod naming and the executing-pod surface
A run's executing pod name is tracked so the UI can show where a run is running.
PodNameRegistryis an in-memory map from run id → bound pod name. It is populated by the Kubernetes sandbox executor once aSandboxClaimreports itsReadyconditionTrue, and the entry is removed when the claim is deleted (e.g. on run cleanup or release).- The registry is consumed in two places:
- the system runtime endpoint (
GET /api/system/runtime) returns{ kubernetes, podName }, wherepodNameis the API/host pod name when running inside Kubernetes — the global fallback; and - the run graph endpoint (
GET /api/runs/{id}/graph) populates anexecutionPodNamefield on each node from the registry, so a per-run/per-node pod name overrides the global fallback as the pod-per-run rollout begins carrying the correct per-pod value automatically.
- the system runtime endpoint (
GET /api/system/runtimereports the host/API pod, not a fallback execution attribution for coordinator children. Child and workflow nodes use topologyexecutionPodNameor null; an unbound child must not be labelled as executing on the API pod.
| Field | Source | Meaning |
|---|---|---|
kubernetes | GET /api/system/runtime | Whether the backend is running inside Kubernetes; gates whether any pod pill is shown. |
podName | API/host pod identity, not fallback attribution for a Coordinator child. | |
executionPodName | Authoritative bound execution pod for this run/node, or null. |
The same
PodNameRegistryalso lets preview/port-forward tooling locate a run's pod. That preview path is documented in the Sandbox deep dive and, for its API surface, in Sandbox preview port-forward below.
Sandbox preview port-forward (Feature 017)
Dedicated pages: this feature now has its own Reference, User Guide, and Deep Dive. The summary below stays here for context within the sandbox-pods surface.
With Sandbox:Preview:Enabled=true, the API creates Gateway-direct HTTPS routing to the run's sandbox and returns preview_url and keepalive_url. Browser traffic bypasses the API. See the sandbox browser preview reference for authorization, approval, publication, keepalive and stop semantics.
Viewer can list previews; Contributor/Owner can start, retry, keep alive or stop them. Legacy non-project runs retain submitting-principal ownership; trusted internal agent callbacks are explicitly scoped exceptions.
When preview creation is disabled, the operator start route uses the legacy kubectl implementation. This is not an automatic fallback after Gateway publication failure.
| Disabled-preview fallback | Scope |
|---|---|
| Bound pod required | Uses the process's PodNameRegistry; a local executor without a Kubernetes pod has nothing to forward. |
kubectl port-forward --address 127.0.0.1 pod/{pod} :{targetPort} -n {namespace} | Binds loopback on the API host. A remote browser's localhost is not that host. |
local_port | No public preview URL; arrange an appropriate local/operator connection separately. |
| Ports / caps | 1-65535; defaults 3 per run, 20 per service process. Gateway has its own configured allowed range. |
| Lifetime | Process-local, no persisted route annotations or session TTL; stop, exit, disposal or run/pod cleanup ends it. |
Security properties
| Property | Pod-per-run guarantee |
|---|---|
| Execution isolation | Each run's agent turn, tools, shell, and file ops run in the run's own Kata-isolated pod (kata-vm-isolation), not a shared process. |
| Control-plane isolation | The orchestration graph, HITL decisions, and run record stay in the worker; a compromised pod cannot alter what happens next. |
| Capability boundary | The API's GitHubCapabilityBroker fences immutable purpose-bound snapshots before and after redemption. AgentHost receives the bounded capability through /configure, without ambient Key Vault/filesystem user-secret lookup. |
| A2A turn auth | message:stream requires Authorization: Bearer {per-run token}. The token is delivered only to the claimed AgentHost pod via /configure and removed from the registry when the pod is released. |
| GitHub token exposure | Brokered by the API for the configured run owner only and delivered in the one-time /configure call, then cached in memory for the pod lifetime; the sandbox identity has no Key Vault access (issue #471), and no CSI user-token file or shared workspace copy exists. |
| Egress | Default-deny with explicit API/MCP/DNS paths and public HTTPS excluding private/link-local ranges; not a per-run Git-host-only allowlist. No direct PostgreSQL access. |
| At rest / past run | Token material does not persist past the pod lifetime; no per-run Secret/SPC is created, and the bearer token is no longer written to SandboxClaim.spec.env in etcd. |
| Reversibility | Change Sandbox:AgentExecutionMode to in-api through the normal configuration rollout; startup DI wiring is not hot reloaded. |
Related reference
- Sandbox setup — operator install/config of the sandbox backends.
- API reference — the endpoints surfaced above.
- A2A reference — the
-previewtransport (experimental) that carries agent turns. - Sandbox pod execution deep dive — the reasoning.
- Sandbox pod execution experience — the user/operator view.
- Sandbox browser preview — preview routes (start/keepalive/stop) that expose a pod-internal server over a public HTTPS reverse proxy.
- Tool Approval SSE Contract — public approval outcomes and coordinator-to-child routing.
Diagram details and constraints
| Element | Contract |
|---|---|
| title | Preview readiness follows the public path |
| takeaway | Provision the route, then probe its exact HTTPS URL; object creation alone is not ready. |
| group-title0 | CONTROL: PROVISION + PROBE |
| group-title1 | GATEWAY DATA PATH |
| Preview API | Preview API |
| Preview API | Resolve bound SandboxClaim |
| Preview API | Patch run selector on pod |
| Preview API | Create Service + HTTPRoute |
| Preview API | State from cluster, not cache |
| Publication probe | Publication probe |
| Publication probe | Exact generated HTTPS URL |
| Publication probe | Wait for managed DNS |
| Publication probe | Check Gateway + application |
| Publication probe | Only then return ready |
| Browser preview | Browser preview |
| Browser preview | Open the returned URL |
| Browser preview | Run-scoped capability host |
| Browser preview | Keepalive via API |
| Browser preview | Iframe: no-referrer |
| Preview Gateway | Preview Gateway |
| Preview Gateway | Separate shared Gateway |
| Preview Gateway | HTTPS host match |
| Preview Gateway | HTTPRoute selects Service |
| Preview Gateway | Not API port-forward |
| ClusterIP Service | ClusterIP Service |
| ClusterIP Service | Per-preview target selector |
| ClusterIP Service | Service :80 → public port |
| ClusterIP Service | Routes to bound sandbox pod |
| ClusterIP Service | Allowed ports 3000–9000 |
| Sandbox preview app | Sandbox preview app |
| Sandbox preview app | AgentHost pod-local path |
| Sandbox preview app | Live preview: TCP forwarder |
| Sandbox preview app | 0.0.0.0 → loopback app |
| Sandbox preview app | Manual: chosen target port |
| relation-0 | 1 after create |
| relation-1 | 2 ready URL |
| relation-2 | 3 HTTPS probe |
| relation-3 | 4 HTTPS |
| relation-4 | 5 route |
| relation-5 | 6 public port |
| assurance | No API → pod TCP readiness probe. Publication failure rolls back; DNS convergence has a bounded retry window. |
| assurance-0-label | Public readiness |
| assurance-0-fact | Probe the exact generated HTTPS URL. |
| assurance-0-source | SandboxPreviewService.cs |
| assurance-1-label | Rollback on failure |
| assurance-1-fact | Unpublish failed preview resources. |
| assurance-1-source | SandboxPreviewPublicationTests.cs |
| assurance-2-label | Separate ingress |
| assurance-2-fact | DNS managed externally, not by API. |
| assurance-2-source | gateway-preview.yaml |
| n0 | Patch run selector on pod; Create Service + HTTPRoute |
| n1 | Wait for managed DNS; Check Gateway + application |
| n2 | Run-scoped capability host; Keepalive via API |
| n3 | HTTPS host match; HTTPRoute selects Service |
| n4 | Service :80 → public port; Routes to bound sandbox pod |
| n5 | Live preview: TCP forwarder; 0.0.0.0 → loopback app |
| groups | CONTROL: PROVISION + PROBE; GATEWAY DATA PATH |
