Operations Guide
This guide covers the day-to-day operational procedures for running and releasing Agentweaver in production on AKS.
For provisioning, image builds, deployment, and verification, use the root package scripts (pnpm run preferred; npm run is equivalent). The AKS deployment runbook defines the canonical sequence and individual-step instructions.
Release Process
Agentweaver uses Changesets and the protected dev → release/vX.Y.Z → main flow. release:prepare generates the version mirrors and changelog on the release branch.
Before preparing or publishing, deploy the exact committed candidate with npm run azure:deploy-from-commit -- <candidate-sha>. Run representative integration and feature-specific API/UI E2E acceptance against that deployment and record passing results bound to the SHA. A changed candidate must be redeployed and retested.
Only after that acceptance passes, prepare and promote the release as described in RELEASING.md. From the exact promoted main SHA, publication and deployment are independent:
npm run release:publish
npm run azure:deploy-from-release -- vX.Y.Zrelease:publish creates the annotated tag, waits for GHCR images, and creates the GitHub Release. azure:deploy-from-release requires that existing published tag, imports or rebuilds its images, deploys them, and verifies the live environment. By default it imports the images already published for that tag by .github/workflows/publish-images.yml instead of rebuilding them from source — this is the fastest way to deploy an already-tagged release and never touches cluster/ACR/Postgres/identity infrastructure. Add --image-source acr-build to rebuild the images from source into ACR instead:
npm run azure:deploy-from-release -- vX.Y.Z --image-source acr-buildFor the normal first shipment to the default environment:
npm run azure:releaseThis composes publication and deployment. It never calculates or commits a version. See RELEASING.md for preparation and recovery, and the Agentweaver changelog skill for the full fragment lifecycle, recovery commands, and changelog/release-notes rules.
Image tags
| Tag | Meaning |
|---|---|
vX.Y.Z | Immutable semver release tag |
<git-sha> | Short SHA from ad-hoc builds (CI / development) |
Note:
latestandlatest-releaseare rejected by the variable scripts. Select a release tag or an immutable commit tag.
Verifying a deployed version
Check the running image tag
kubectl get deployment agentweaver-api \
--namespace agentweaver \
--output jsonpath='{.spec.template.spec.containers[0].image}'Check OCI image labels (version + commit SHA)
az acr repository show-tags \
--name agentweaverregistry \
--repository agentweaver-api \
--orderby time_desc \
--top 5Or inspect labels on the image:
# Pull the manifest (no local pull needed)
az acr manifest show \
agentweaverregistry.azurecr.io/agentweaver-api:v0.6.1Each image is built with the following OCI labels:
| Label | Value |
|---|---|
org.opencontainers.image.version | Semver tag (e.g. v0.6.1) |
org.opencontainers.image.revision | Full git commit SHA |
Rolling back a release
Before rollback, verify that the previous version remains compatible with the current database schema, credential storage, and backing configuration. A retained image tag alone does not establish rollback safety. In particular, the Fleet cutover boundary does not permit recreating deleted legacy credentials after irreversible cleanup. When those checks permit rollback, deploy the exact published release:
npm run azure:deploy-from-release -- v0.6.0All previous semver tags remain in ACR and are not deleted by the release process.
Manual image builds (development)
To build and push images without cutting a release (e.g. for a staging environment), use azure:deploy-from-local, which builds using the current git SHA as the tag and then redeploys:
npm run azure:deploy-from-localTo deploy another committed ref without switching the current checkout:
npm run azure:deploy-from-commit -- <sha-or-ref>Observability notes
- Token and AIC usage data now lives in Application Insights / Azure Monitor, not in the application database.
- The project dashboard throughput chart and agent leaderboard read from
GET /api/projects/{id}/metrics, which proxies App Insights KQL. - Configure
APPLICATIONINSIGHTS_CONNECTION_STRINGand a Log Analytics workspace id (APPLICATIONINSIGHTS_WORKSPACE_IDorApplicationInsights:WorkspaceId) unless your connection string already embedsWorkspaceId. - If App Insights is not configured, or no workspace id can be resolved, the metrics endpoint returns empty arrays so the dashboard degrades gracefully.
Cluster topology details
The Cluster page topology cards open an operator detail panel instead of repeating the card text. Runtime and workload details are sourced from the bounded GET /api/diagnostics/cluster/topology envelope, which allow-lists concise Kubernetes fields rather than exposing raw manifests or cluster credentials to the browser.
Use the panel to copy pod, claim, run, deployment, warm-pool, sandbox, and template identifiers during triage. Healthy snapshots stay quiet; unhealthy pods, short readiness, and non-zero restarts are sorted first and called out. Each panel shows the topology snapshot's Last updated time and names any partial layer read (for example, runtime detail timeout) instead of falling back to a bare count.
Sandbox details show the runtime class and isolation backend. In AKS, AgentHost sandboxes normally run with kata-vm-isolation and the agentweaver-exec sidecar on the Kata node pool, while non-sandbox control-plane workloads run with the default runc runtime.
API restart recovery and health probes
API and worker restart recovery run after their listeners start. A shared Postgres advisory lock serializes sweeps across both roles. A successful API or worker leader holds it until that process stops, so no other replica can repeat a completed startup sweep over newly created runs. Followers retry acquiring leadership without a fixed attempt limit; if the leader exits or dies, either role can take over recovery. A failed or timed-out leader sweep releases the lock and retries up to three total sweep attempts per process. After the third failure, that process stops startup recovery and logs exhaustion; another replica can still acquire the lock. The advisory lock does not authorize mutations: durable run leases, coordinator plan claims, and child-dispatch reservations fence work across API and worker roles. Healthy child work on a surviving replica remains associated with its existing run identity. An expired coordinator child is restarted under that same child run ID after claiming its execution lease; a fresh lease held by another replica is skipped. When a coordinator terminates, its assemble-ready child and revision sandboxes are released after the final Scribe turn. A stopped coordinator releases them after the stop settles. On API startup and each coordinator heartbeat, a bounded page of terminal coordinators is revisited to recover claims left by interrupted cleanup. A child with a current durable execution lease or pending review is preserved; a terminal-parent child stuck Pending/InProgress without a live lease is reclaimed. A live preview retains its pod until preview expiry. Release reads one claim snapshot and uses Kubernetes UID/resourceVersion delete preconditions alongside the holder and lifecycle generation, so a replacement claim cannot be deleted by an older cleanup. Failures are logged as Terminal child cleanup warnings for retry rather than changing the coordinator's outcome. If a new revision stays pending on Kata, inspect these warnings and the child claim inventory before considering cluster capacity changes. An unbound AgentHost claim emits sandbox.provisioning_pending with the pod's PodScheduled=False reason when available; the coordinator displays that reason. If a pod cannot schedule before Sandbox:Kubernetes:AgentHostProvisioningTimeoutSeconds (default 600 seconds), the launch fails with the latest scheduling diagnosis and releases its claim. Check the pending pod's scheduler condition and node-pool autoscaler before retrying. Do not stop unrelated users' previews to free capacity. For a current preview, inspect GET /api/runs/{id} sandbox.current_binding: verified identifies the configured claim UID, Pod UID, namespace, generation, attempt, and source tree; unavailable or conflict includes a reason and must not be replaced with the historical sandbox.pod_name. A released execution lease and a retained child-owned preview are distinct lifecycle facts. A child's claim, Pod, and preview session do not attest its coordinator parent's claim or automatic preview. Older claims lacking the post-configure attestation remain explicitly unavailable. After a durable agent.turn.end, coordinator observation first waits Coordinator:PostTurnFinalizationGraceSeconds (default 10 seconds, clamped to 0.1–30 seconds) for assemble-ready or another terminal event. If the recovered child still owns an unexpired execution lease, it rechecks durable completion while that lease remains active, up to Coordinator:PostTurnFinalizationMaxWaitSeconds (default five minutes, clamped between the grace and ten minutes). An absent/expired lease or exhausted cap restores normal stall recovery; neither setting changes lease fencing. Inspect the child's execution lease and terminal run events before increasing the cap. Both API replicas can answer /api/ping without waiting for a sweep. A sweep has a five-minute deadline; followers and failed leader sweeps retry after 30 seconds, but a successful leader does not resweep. Look for Startup recovery sweep started, completed, exceeded, failed, or exhausted in API logs when diagnosing a restart. Readiness reflects workspace availability and successful initial static OAuth client reconciliation, not completion of the recovery backlog: operators should check the sweep log before assuming every interrupted run has been re-armed. /api/health, /healthz/workspace, and /oauth/* return 503 until the initial static OAuth client reconciliation succeeds; /api/ping stays responsive throughout. Failed reconciliations are logged and retried every five seconds after a 30-second attempt deadline without terminating the host. Database migrations and the bounded Copilot App registration validation still precede serving traffic.
AgentHost pre-delivery recovery diagnostics
Project agents, Assembly RAI, and Build & Test use warm-pool AgentHost claims. Before the first A2A request, the dispatch is generation-fenced and recovery is bounded:
agenthost_configure_copilot_token_refreshedmeans AgentHost explicitly rejected the configured Copilot credential, the API rotated that exact user/account scope, and Build & Test will recreate the one-time-configured pod once.agenthost_configure_copilot_unauthorizedmeans no different credential could be produced. The failure is not retried; the submitting user must repair GitHub/Copilot authorization.- Readiness, one-time configuration, a missing pod endpoint, or a reaped pod permits one fresh claim only while no model turn has been delivered.
- Exhaustion writes one
agent_host_unavailableterminal outcome withretryable: true. Retry creates a fresh run generation; the failed generation is not left live.
Logs include RunId, pod name, reason, recovery action, and bounded attempt counts. Token values are never logged, and the persisted terminal payload contains only the canonical error, retryability, and opaque dispatch id. Authorization/provider failures remain their specific typed errors rather than being converted to availability failures.
Once message:stream delivery starts, Agentweaver does not retry inside RemoteAgentProxy: request acceptance is uncertain, so replay could execute the same model turn twice. Post-acceptance transport and turn failures continue through the structured turn-failure reasons below. Operator Assistant conversations remain resumable; an AgentHost failure is recorded as a turn error and does not terminalize the conversation.
Diagnosing agent turn infrastructure failures
Agent turns executed through AgentHost use structured terminal reasons. agent_turn_internal_error with retryable: true is the fallback for an unstructured run.failed, an unsupported or unset A2A event, or a pod bridge turn that throws before emitting a structured terminal. Other A2A exceptions use a2a_transport_failure, whose retryability follows the transport error. A clean stream that ends without agent.turn.end uses retryable agent_host_turn_incomplete. In the collective Build & Test stage, the assembly reason prefixes the applicable reason with build_test_infra_, for example build_test_infra_agent_host_turn_incomplete.
This fallback does not hide more specific outcomes. Caller cancellation remains cancellation, and typed timeouts or failures retain their original error code and retryability. A retryable terminal means the workflow may safely consider a bounded retry or redispatch; it never converts the interrupted turn into a success.
For collective assembly, a typed retryable RAI provider or infrastructure failure receives one gate-only retry. The coordinator re-verifies the persisted aggregate tree hash and diff before retrying RAI; it does not rebuild integration, redispatch children, replace artifacts, or rerun completed Build & Test evidence. Content verdicts and revision feedback remain on the normal steering and human-review paths.
To investigate:
- Read
GET /api/runs/{id}/terminal-diagnosticorrun_failure_diagnostic. - Start with
observed_facts, then reviewsupported_interpretations. Do not treat a nearby or repeated tool error as causal unless the terminal evidence directly references the same call or gate. - Check
evidence_sourcesandcompleteness. Missing telemetry does not erase durable terminal evidence, but partial or unavailable sources are not a healthy result. - If
denial_gateis present, repair the named authorization/configuration gate before retrying. A pending human approval is waiting, not denial. - Follow the structured
next_actions; they describe preconditions and effects but do not mutate the run, policy, or authorization.
Diagnostics in the run event are deliberately bounded, flattened to one line, and credential-redacted. They are safe context for triage, not a replacement for restricted server-side logs.
Cluster inventory collection uses the same explicit absence rule. Each inventory_sources entry reports available, no_resources, forbidden, timeout, unsupported, malformed, or collection_error. Only no_resources means collection completed successfully and found nothing. The Cluster page warns when any source is incomplete instead of presenting an unavailable inventory as an empty healthy one.
Related scripts
| Command | Purpose |
|---|---|
npm run azure:release | Full semver release (see above) |
npm run release:publish | Create the annotated tag and GitHub Release without deploying |
npm run azure:deploy-from-release -- vX.Y.Z [--image-source acr-build] | Deploy an existing published release (import already-published GHCR images by default, or rebuild from source) |
npm run azure:provision-infra | Provision/redeploy AKS, identity, monitoring, OAuth signing key, and PostgreSQL |
npm run azure:deploy-from-local | Build, push, and verify images in ACR, then redeploy and cycle the warm pool |
npm run azure:deploy-from-commit -- <sha-or-ref> | Deploy an arbitrary exact commit through a temporary detached worktree |
npm run azure:verify | Verify the current deployment |
Use pnpm run in place of npm run if pnpm is your selected package runner. The runbook's individual-step section shows how to rerun one step.
Observability
Agentweaver ships with end-to-end telemetry using Azure Monitor OpenTelemetry Distro (Application Insights) and AKS Managed Prometheus.
Inspecting a transaction trace
Open a project, select Observability → Traces, then choose Preview trace for a coordinator run. The trace detail includes a timeline, span attributes, and persisted run events. Trace spans load in chronological pages; choose Load more spans until no more spans are available to inspect the complete trace. The opaque continuation keeps already loaded spans, selection, and tree state intact, and a failed page can be retried without reloading the whole trace. Tool spans carry bounded, redacted input/output previews in Application Insights, and persisted events are still loaded when the Events tab or a tool span needs additional context. It shows only trace data returned by Application Insights and the persisted run-event API. In particular, it shows the trace session ID only when the runtime emitted one, and it does not invent event timestamps when a legacy persisted event has no recorded time. The attributes pane is a fixed, safe schema rather than a dump of custom dimensions: it includes operational identity, model, provider, policy, sandbox, usage, status, and tool-payload capture state fields, but never prompts, credentials, raw tokens, secrets, unbounded output, or arbitrary tool payloads. For coordinator runs, child-run spans are grouped below the child agent that executed them using the persisted parent-run relationship; the original distributed trace parent remains available in the span data. See Transaction traces for the span and tool-call details.
If the Application Insights workspace is unavailable or slow, trace retrieval uses a 30-second server-side query budget by default, separately from the three-second dashboard-metrics budget. Set Metrics__AppInsights__TraceQueryTimeoutSeconds or APPINSIGHTS_TRACE_QUERY_TIMEOUT_SECONDS (1–60 seconds) to tune that trace budget. The trace panel automatically retries short-lived dependency failures with backoff, but surfaces a bounded query timeout without repeating the full long-running request. While a retry remains, the panel stays in its normal loading state instead of showing a failure banner. The API coalesces concurrent requests for the same run and cursor page into one bounded workspace query, and trace reads are not short-circuited by an unrelated dashboard-metrics cooldown. A recently retrieved page may be shown while the source recovers and is explicitly labeled as such; an unavailable source with no safe cached page is not presented as proof that the run has no trace data. Cursor paging remains incremental, so retry or Load more spans only requests the needed page. If the automatic attempts are exhausted, use Retry to start a fresh bounded trace load; platform operators can use the API log's query context and failure type to investigate workspace credentials, RBAC, and availability without logging KQL payloads.
Provisioning monitoring resources
Monitoring is provisioned as part of npm run azure:provision-infra. To rerun only that step, see the runbook's individual-step section (scripts/azure/steps/15-provision-monitoring.mjs).
This creates:
- A Log Analytics workspace (
agentweaver-logs) - A workspace-based Application Insights resource (
agentweaver-insights) — workspace-based is required for the Agents (Preview) view - Stores the connection string as
appinsights-connection-stringin Key Vault - Enables AKS Managed Prometheus on the cluster
Finding the Application Insights resource
- Open the Azure Portal
- Navigate to your resource group (
agentweaver-rgby default) - Select the Application Insights resource named
agentweaver-insights
Using the Agents (Preview) view
The Agents (Preview) view in Application Insights shows GenAI-specific telemetry including agent runs, token usage, and model calls.
- In the Application Insights resource, select Agents (Preview) from the left menu
- Use the time range picker to scope your investigation
- Filter by agent using the
gen_ai.agent.nameattribute — this maps to the configured agent name in the squad definition (e.g.morpheus,seraph)
Key span attributes emitted by Agentweaver:
| Attribute | Description |
|---|---|
gen_ai.agent.name | Squad agent name |
gen_ai.agent.id | Agent identifier |
gen_ai.usage.input_tokens | Prompt tokens consumed |
gen_ai.usage.output_tokens | Completion tokens produced |
gen_ai.request.model | Model deployment name |
gen_ai.operation.name | chat or execute_tool |
Querying a specific run in Application Insights Search
To find all telemetry for a single run by its RunId:
- In Application Insights, select Search (or Transaction search)
- Enter the RunId (e.g.
run_abc123) in the search box - Alternatively, use Logs with a KQL query:
traces
| where customDimensions["RunId"] == "run_abc123"
| order by timestamp ascOr to see all token usage for a run:
customMetrics
| where name == "agentweaver.token.usage"
| where customDimensions["run_id"] == "run_abc123"
| summarize totalTokens = sum(value) by tostring(customDimensions["agent_name"])AKS Managed Prometheus metrics
Business metrics emitted by AgentWeaverMetrics are exported to the AKS Managed Prometheus workspace:
| Metric | Type | Description |
|---|---|---|
agentweaver_token_usage_total | Counter | Token usage by agent and model |
agentweaver_run_duration | Histogram | Run duration in milliseconds |
agentweaver_run_errors_total | Counter | Run errors by type |
agentweaver_run_active | UpDownCounter | Currently active runs |
agentweaver_run_queued | Gauge | Active-project Ready backlog tasks awaiting coordinator pickup (backlog_tasks.state='ready' AND run_id IS NULL), sampled every 15s. Legacy name retained; aggregate with max, not sum, because every replica exports the same global snapshot. |
To query in Azure Managed Grafana (linked to the Prometheus workspace), use standard PromQL:
rate(agentweaver_token_usage_total[5m])Worker autoscaling (queue depth vs. resource utilization)
k8s/base/worker-hpa.yaml scales agentweaver-worker between 2 and 3 replicas using CPU utilization at 70% and memory utilization at 80%. Neither metric is the Ready queue depth; resource utilization alone is an imperfect backlog proxy.
The agentweaver_run_queued gauge above (issue #108) exists specifically to provide a real queue-depth signal for this HPA. In the current system that signal is notruns.status='pending' — backlog pickup creates coordinator runs directly as in_progress. The durable queue is the set of active-project backlog tasks still in Ready with no bound run_id yet, which is what the gauge now publishes every 15 seconds. Because each replica exports the same shared-store total, Prometheus/KEDA queries must use max(agentweaver_run_queued) (or the equivalent single-series selector), not sum(...).
The HPA itself has not yet been switched over — wiring an external metric type into a plain HorizontalPodAutoscaler requires a Kubernetes External Metrics API adapter capable of serving Azure Monitor managed-Prometheus-backed queries, and no such adapter is currently provisioned in scripts/azure/. The two realistic paths forward (tracked against #108 — see decisions/inbox/niobe-108-hpa-investigation.md for the full analysis) are:
- KEDA with a Prometheus scaler (Microsoft's supported pattern for scaling on Azure Monitor managed Prometheus metrics) — query
agentweaver_run_queuedvia the workspace's Prometheus query endpoint usingmax(...), notsum(...). - Provision a
k8s-prometheus-adapter-style External Metrics API adapter and wireworker-hpa.yamlwith atype: Externalmetric block pointing at it.
Until one of these is chosen and the supporting cluster component is provisioned, the worker continues to scale on CPU (with the gauge available for manual/Grafana-based capacity monitoring in the meantime).
