Skip to content

Cluster page ​

The Cluster page gives operators a real-time view of the Kubernetes cluster backing the Agentweaver AKS deployment: sandbox capacity, Kubernetes health checks, and any legacy subtasks recorded as waiting for capacity.

Kubernetes owns scheduling (issue #217)

The platform no longer pre-gates on quota. It submits the SandboxClaim and waits for Kubernetes to schedule and bind the pod, so a Pending pod is an expected transient state, not a failure. The namespace ResourceQuota no longer caps CPU/memory (only object counts). The Pending capacity KPI, the Waiting for capacity badge, and the pending-capacity table are back-compat surfaces that only reflect historical runs.

It is available under the Cluster nav item in the SYSTEM section of the project left rail. Route: /projects/:projectId/cluster. The page auto-refreshes every 30 seconds by default; you can toggle Auto-refresh off when you want to inspect a static snapshot.

Read the health summary and Sandbox claims together: a claim waiting to bind is different from an orphaned pod. The live Resource topology provides the current inventory; no static capture is needed to infer which resources exist in your deployment.

When to use the Cluster page ​

Open the Cluster page when:

  • a coordinator run shows subtasks in ⏳ Waiting for capacity (amber badge in the topology graph);
  • runs are dispatching slowly and you suspect pod scheduling or node-pool autoscaling delays;
  • you want to confirm that all Kubernetes API components are reachable;
  • orphaned pods are accumulating (the reaper has not swept them yet);
  • after a deployment or scaling event, to confirm the cluster is healthy.

KPI cards ​

The KPI cards at the top of the page summarize the cluster signals the current UI exposes:

CardWhat it shows
Orphaned podsAgent pods reported as orphaned by cluster diagnostics.
Pending capacityLegacy. Subtasks recorded in the historical PendingCapacity status; empty for new runs (Kubernetes now owns scheduling).
Checks healthyHealthy checks divided by all reported cluster checks.
Warm pool readyReady vs. desired warm sandbox replicas when warm-pool data is available.

Below the KPIs, the page shows Health checks, Sandbox claims, Orphaned agent pods when present, Pending capacity, and Warm pools.

Component health table ​

Cluster checks run concurrently each time the page loads:

CheckWhat it testsTypical failure cause
PostgresConnectivity to the Postgres databaseNetwork policy, password rotation
Azure Key VaultCSI delivery of the required mcp-api-keyManaged identity misconfiguration, network policy, or a missing API authentication secret
Agent pod quotaEffective admission headroom from the enforced pods and SandboxClaim object quotas. Healthy means plenty of room remains, warning means only a handful of starts remain, and critical means no new AgentHost can be admitted.Namespace object-quota exhaustion
Warm poolReadiness of the configured AgentHost warm pool (default agentweaver-agent-host, checked-in target 2)Ready replicas below target, missing pool or template
Kubernetes APIKubernetes API server reachabilityIn-cluster network policy, apiserver overload

Each check shows:

  • A status badge such as healthy, warning, degraded, or critical.
  • A detail message (visible on warn/fail) explaining the specific failure.
  • The duration the check took in milliseconds.

All five checks have a 5-second individual timeout. The guarded timeout result is unknown with the detail "check timed out"; it is not a healthy result.

If the Key Vault row shows critical: secret 'mcp-api-key' not found, restore the required API authentication secret with npm run azure:provision-infra before redeploying.

Active execution placement ​

The current page does not render a separate active-agent-pods table or CPU/memory quota bars. Use Sandbox claims and the Runtime topology layer for live execution resources, and the run's per-node pod indicator for recorded placement. Preview retention can keep resources alive beyond a run; do not infer that every retained pod is orphaned merely because the run is terminal.

Orphaned agent pods table ​

Lists pods that are running but have no matching active run. These will be terminated on the next reaper sweep (default: every ~2 minutes).

If orphaned pods are not being cleaned up, check:

  • That the heartbeat is enabled and ticking (see the Heartbeat page).
  • That Coordinator:ReaperIntervalTicks is not set to an unusually large value.

Pending-capacity runs table ​

Legacy / back-compat. This table renders historical PendingCapacity records from the removed pre-gating loop. A zero count does not prove current pods scheduled immediately. For a live run whose pod is still being scheduled, inspect its claim and sandbox.provisioning_pending heartbeats instead.

Lists coordinator subtasks recorded in the historical PendingCapacity status:

ColumnMeaning
Coordinator run IDThe parent coordinator run
SubtaskThe subtask that was waiting
Pending sinceWhen the subtask entered PendingCapacity
Retry countHow many dispatch attempts were made under the removed park/retry loop

Warm pools table ​

Lists every SandboxWarmPool CRD object in the namespace. Each row represents one pool:

ColumnMeaning
NameKubernetes name of the SandboxWarmPool object
DesiredTarget number of pre-warmed sandboxes declared in the pool spec
ReadySandboxes currently ready to accept a claim
AvailableSandboxes that are ready and not yet claimed by a run
Statushealthy when ready meets a positive desired count; warning when some are ready but below target; critical when none are ready

A pool below its target can increase claim-binding latency. Warm-pool availability is not an application reservation gate: the controller and Kubernetes still own provisioning and scheduling.

Resource topology ​

The Resource topology groups the live Kubernetes inventory by Agentweaver function. The default Runtime layer shows session and agent execution plus the full sandbox lifecycle: SandboxTemplate → SandboxWarmPool → SandboxClaim → Sandbox. Each item is a first-class card with aggregate live counts, rather than a tall card stack of pod names.

Use the layer controls to add the few related functions you need:

LayerFunctional view
RuntimeAgent execution, sandbox templates, pools, claims, and sandboxes.
NetworkingPublic entry and gateway, grouped Agentweaver service targets, and public-ingress NetworkPolicy guardrails.
WorkloadsAgentweaver deployment workloads and their pod workloads. ReplicaSets are not shown.
StorageApplication persistence and the shared sandbox/session artifact workspace when present.
AutoscalingThe capacity policies that scale Agentweaver workloads.
AvailabilityAvailable disruption-protection signals.

The traffic flow is intentionally small: public entry and gateway → Agentweaver service targets → control plane. Green edges mean an explicit gateway-ingress NetworkPolicy is present. The Network policy guardrails card also shows a concise allow/deny badge: explicit gateway allows and the default policy that blocks other inbound traffic. Select the card to inspect each reported policy's selector, direction, and effect. Egress, internal-only, and sandbox-only rules are not rendered into the traffic flow.

The storage view exposes only meaningful persistence functions: Application persistence and Sandbox & session artifacts. It does not expand PVCs, volumes, mounts, persistent volumes, or storage classes into the graph. Edges show the control plane persisting application state and agent execution writing artifacts.

All cards use checked-in IconCloud Azure assets. Select a card to inspect its bounded function summary. The Cluster page never displays full manifests, Secret data, tokens, internal addresses, or container environment values.

Sandbox claims table ​

Lists all SandboxClaim CRD objects in the namespace:

ColumnMeaning
NameKubernetes name of the SandboxClaim object
Phasebound (assigned to a sandbox), pending (waiting for one), or unknown
ReadyWhether the claimed sandbox is ready
RunThe run ID that created this claim, linking to an orchestration detail when present
Bound sandboxThe Sandbox object this claim is bound to; blank when still pending
Warm poolThe pool the bound sandbox came from; blank for ad-hoc claims
AgeHow long the claim has existed

404 fallback ​

When the API is not deployed on AKS, or the cluster diagnostics endpoint is unavailable, the page displays a message indicating that cluster diagnostics are not available for this deployment. No other page functionality is affected.