Cluster page
The Cluster page gives operators a real-time view of the Kubernetes cluster backing the Agentweaver AKS deployment: sandbox capacity, Kubernetes health checks, and any legacy subtasks recorded as waiting for capacity.
Kubernetes owns scheduling (issue #217)
The platform no longer pre-gates on quota. It submits the SandboxClaim and waits for Kubernetes to schedule and bind the pod, so a Pending pod is an expected transient state, not a failure. The namespace ResourceQuota no longer caps CPU/memory (only object counts). The Pending capacity KPI, the Waiting for capacity badge, and the pending-capacity table are back-compat surfaces that only reflect historical runs.
It is available under the Cluster nav item in the SYSTEM section of the project left rail. Route: /projects/:projectId/cluster. The page auto-refreshes every 30 seconds by default; you can toggle Auto-refresh off when you want to inspect a static snapshot.
Read the health summary and Sandbox claims together: a claim waiting to bind is different from an orphaned pod. The live Resource topology provides the current inventory; no static capture is needed to infer which resources exist in your deployment.
When to use the Cluster page
Open the Cluster page when:
- a coordinator run shows subtasks in ⏳ Waiting for capacity (amber badge in the topology graph);
- runs are dispatching slowly and you suspect pod scheduling or node-pool autoscaling delays;
- you want to confirm that all Kubernetes API components are reachable;
- orphaned pods are accumulating (the reaper has not swept them yet);
- after a deployment or scaling event, to confirm the cluster is healthy.
KPI cards
The KPI cards at the top of the page summarize the cluster signals the current UI exposes:
| Card | What it shows |
|---|---|
| Orphaned pods | Agent pods reported as orphaned by cluster diagnostics. |
| Pending capacity | Legacy. Subtasks recorded in the historical PendingCapacity status; empty for new runs (Kubernetes now owns scheduling). |
| Checks healthy | Healthy checks divided by all reported cluster checks. |
| Warm pool ready | Ready vs. desired warm sandbox replicas when warm-pool data is available. |
Below the KPIs, the page shows Health checks, Sandbox claims, Orphaned agent pods when present, Pending capacity, and Warm pools.
Component health table
Cluster checks run concurrently each time the page loads:
| Check | What it tests | Typical failure cause |
|---|---|---|
| Postgres | Connectivity to the Postgres database | Network policy, password rotation |
| Azure Key Vault | CSI delivery of the required mcp-api-key | Managed identity misconfiguration, network policy, or a missing API authentication secret |
| Agent pod quota | Effective admission headroom from the enforced pods and SandboxClaim object quotas. Healthy means plenty of room remains, warning means only a handful of starts remain, and critical means no new AgentHost can be admitted. | Namespace object-quota exhaustion |
| Warm pool | Readiness of the configured AgentHost warm pool (default agentweaver-agent-host, checked-in target 2) | Ready replicas below target, missing pool or template |
| Kubernetes API | Kubernetes API server reachability | In-cluster network policy, apiserver overload |
Each check shows:
- A status badge such as
healthy,warning,degraded, orcritical. - A detail message (visible on warn/fail) explaining the specific failure.
- The duration the check took in milliseconds.
All five checks have a 5-second individual timeout. The guarded timeout result is unknown with the detail "check timed out"; it is not a healthy result.
If the Key Vault row shows critical: secret 'mcp-api-key' not found, restore the required API authentication secret with npm run azure:provision-infra before redeploying.
Active execution placement
The current page does not render a separate active-agent-pods table or CPU/memory quota bars. Use Sandbox claims and the Runtime topology layer for live execution resources, and the run's per-node pod indicator for recorded placement. Preview retention can keep resources alive beyond a run; do not infer that every retained pod is orphaned merely because the run is terminal.
Orphaned agent pods table
Lists pods that are running but have no matching active run. These will be terminated on the next reaper sweep (default: every ~2 minutes).
If orphaned pods are not being cleaned up, check:
- That the heartbeat is enabled and ticking (see the Heartbeat page).
- That
Coordinator:ReaperIntervalTicksis not set to an unusually large value.
Pending-capacity runs table
Legacy / back-compat. This table renders historical
PendingCapacityrecords from the removed pre-gating loop. A zero count does not prove current pods scheduled immediately. For a live run whose pod is still being scheduled, inspect its claim andsandbox.provisioning_pendingheartbeats instead.
Lists coordinator subtasks recorded in the historical PendingCapacity status:
| Column | Meaning |
|---|---|
| Coordinator run ID | The parent coordinator run |
| Subtask | The subtask that was waiting |
| Pending since | When the subtask entered PendingCapacity |
| Retry count | How many dispatch attempts were made under the removed park/retry loop |
Warm pools table
Lists every SandboxWarmPool CRD object in the namespace. Each row represents one pool:
| Column | Meaning |
|---|---|
| Name | Kubernetes name of the SandboxWarmPool object |
| Desired | Target number of pre-warmed sandboxes declared in the pool spec |
| Ready | Sandboxes currently ready to accept a claim |
| Available | Sandboxes that are ready and not yet claimed by a run |
| Status | healthy when ready meets a positive desired count; warning when some are ready but below target; critical when none are ready |
A pool below its target can increase claim-binding latency. Warm-pool availability is not an application reservation gate: the controller and Kubernetes still own provisioning and scheduling.
Resource topology
The Resource topology groups the live Kubernetes inventory by Agentweaver function. The default Runtime layer shows session and agent execution plus the full sandbox lifecycle: SandboxTemplate → SandboxWarmPool → SandboxClaim → Sandbox. Each item is a first-class card with aggregate live counts, rather than a tall card stack of pod names.
Use the layer controls to add the few related functions you need:
| Layer | Functional view |
|---|---|
| Runtime | Agent execution, sandbox templates, pools, claims, and sandboxes. |
| Networking | Public entry and gateway, grouped Agentweaver service targets, and public-ingress NetworkPolicy guardrails. |
| Workloads | Agentweaver deployment workloads and their pod workloads. ReplicaSets are not shown. |
| Storage | Application persistence and the shared sandbox/session artifact workspace when present. |
| Autoscaling | The capacity policies that scale Agentweaver workloads. |
| Availability | Available disruption-protection signals. |
The traffic flow is intentionally small: public entry and gateway → Agentweaver service targets → control plane. Green edges mean an explicit gateway-ingress NetworkPolicy is present. The Network policy guardrails card also shows a concise allow/deny badge: explicit gateway allows and the default policy that blocks other inbound traffic. Select the card to inspect each reported policy's selector, direction, and effect. Egress, internal-only, and sandbox-only rules are not rendered into the traffic flow.
The storage view exposes only meaningful persistence functions: Application persistence and Sandbox & session artifacts. It does not expand PVCs, volumes, mounts, persistent volumes, or storage classes into the graph. Edges show the control plane persisting application state and agent execution writing artifacts.
All cards use checked-in IconCloud Azure assets. Select a card to inspect its bounded function summary. The Cluster page never displays full manifests, Secret data, tokens, internal addresses, or container environment values.
Sandbox claims table
Lists all SandboxClaim CRD objects in the namespace:
| Column | Meaning |
|---|---|
| Name | Kubernetes name of the SandboxClaim object |
| Phase | bound (assigned to a sandbox), pending (waiting for one), or unknown |
| Ready | Whether the claimed sandbox is ready |
| Run | The run ID that created this claim, linking to an orchestration detail when present |
| Bound sandbox | The Sandbox object this claim is bound to; blank when still pending |
| Warm pool | The pool the bound sandbox came from; blank for ad-hoc claims |
| Age | How long the claim has existed |
404 fallback
When the API is not deployed on AKS, or the cluster diagnostics endpoint is unavailable, the page displays a message indicating that cluster diagnostics are not available for this deployment. No other page functionality is affected.
Related reading
- Operations — all operations surfaces at a glance.
- Cluster diagnostics reference — full API response schema.
- Sandbox pod execution — reaper design and Kubernetes-owned pod admission (
sandbox.provisioning_pending). - Heartbeat — the heartbeat that drives the reaper.
