Skip to content

Operations Guide ​

This guide covers the day-to-day operational procedures for running and releasing Agentweaver in production on AKS.

For provisioning, image builds, deployment, and verification, use the root package scripts (pnpm run preferred; npm run is equivalent). The AKS deployment runbook defines the canonical sequence and individual-step instructions.

Release Process ​

Agentweaver uses Changesets and the protected dev → release/vX.Y.Z → main flow. release:prepare generates the version mirrors and changelog on the release branch.

Before preparing or publishing, deploy the exact committed candidate with npm run azure:deploy-from-commit -- <candidate-sha>. Run representative integration and feature-specific API/UI E2E acceptance against that deployment and record passing results bound to the SHA. A changed candidate must be redeployed and retested.

Only after that acceptance passes, prepare and promote the release as described in RELEASING.md. From the exact promoted main SHA, publication and deployment are independent:

bash
npm run release:publish
npm run azure:deploy-from-release -- vX.Y.Z

release:publish creates the annotated tag, waits for GHCR images, and creates the GitHub Release. azure:deploy-from-release requires that existing published tag, imports or rebuilds its images, deploys them, and verifies the live environment. By default it imports the images already published for that tag by .github/workflows/publish-images.yml instead of rebuilding them from source — this is the fastest way to deploy an already-tagged release and never touches cluster/ACR/Postgres/identity infrastructure. Add --image-source acr-build to rebuild the images from source into ACR instead:

bash
npm run azure:deploy-from-release -- vX.Y.Z --image-source acr-build

For the normal first shipment to the default environment:

bash
npm run azure:release

This composes publication and deployment. It never calculates or commits a version. See RELEASING.md for preparation and recovery, and the Agentweaver changelog skill for the full fragment lifecycle, recovery commands, and changelog/release-notes rules.

Image tags ​

TagMeaning
vX.Y.ZImmutable semver release tag
<git-sha>Short SHA from ad-hoc builds (CI / development)

Note: latest and latest-release are rejected by the variable scripts. Select a release tag or an immutable commit tag.

Verifying a deployed version ​

Check the running image tag ​

bash
kubectl get deployment agentweaver-api \
  --namespace agentweaver \
  --output jsonpath='{.spec.template.spec.containers[0].image}'

Check OCI image labels (version + commit SHA) ​

bash
az acr repository show-tags \
  --name agentweaverregistry \
  --repository agentweaver-api \
  --orderby time_desc \
  --top 5

Or inspect labels on the image:

bash
# Pull the manifest (no local pull needed)
az acr manifest show \
  agentweaverregistry.azurecr.io/agentweaver-api:v0.6.1

Each image is built with the following OCI labels:

LabelValue
org.opencontainers.image.versionSemver tag (e.g. v0.6.1)
org.opencontainers.image.revisionFull git commit SHA

Rolling back a release ​

Before rollback, verify that the previous version remains compatible with the current database schema, credential storage, and backing configuration. A retained image tag alone does not establish rollback safety. In particular, the Fleet cutover boundary does not permit recreating deleted legacy credentials after irreversible cleanup. When those checks permit rollback, deploy the exact published release:

bash
npm run azure:deploy-from-release -- v0.6.0

All previous semver tags remain in ACR and are not deleted by the release process.

Manual image builds (development) ​

To build and push images without cutting a release (e.g. for a staging environment), use azure:deploy-from-local, which builds using the current git SHA as the tag and then redeploys:

bash
npm run azure:deploy-from-local

To deploy another committed ref without switching the current checkout:

bash
npm run azure:deploy-from-commit -- <sha-or-ref>

Observability notes ​

  • Token and AIC usage data now lives in Application Insights / Azure Monitor, not in the application database.
  • The project dashboard throughput chart and agent leaderboard read from GET /api/projects/{id}/metrics, which proxies App Insights KQL.
  • Configure APPLICATIONINSIGHTS_CONNECTION_STRING and a Log Analytics workspace id (APPLICATIONINSIGHTS_WORKSPACE_ID or ApplicationInsights:WorkspaceId) unless your connection string already embeds WorkspaceId.
  • If App Insights is not configured, or no workspace id can be resolved, the metrics endpoint returns empty arrays so the dashboard degrades gracefully.

Cluster topology details ​

The Cluster page topology cards open an operator detail panel instead of repeating the card text. Runtime and workload details are sourced from the bounded GET /api/diagnostics/cluster/topology envelope, which allow-lists concise Kubernetes fields rather than exposing raw manifests or cluster credentials to the browser.

Use the panel to copy pod, claim, run, deployment, warm-pool, sandbox, and template identifiers during triage. Healthy snapshots stay quiet; unhealthy pods, short readiness, and non-zero restarts are sorted first and called out. Each panel shows the topology snapshot's Last updated time and names any partial layer read (for example, runtime detail timeout) instead of falling back to a bare count.

Sandbox details show the runtime class and isolation backend. In AKS, AgentHost sandboxes normally run with kata-vm-isolation and the agentweaver-exec sidecar on the Kata node pool, while non-sandbox control-plane workloads run with the default runc runtime.

API restart recovery and health probes ​

API and worker restart recovery run after their listeners start. A shared Postgres advisory lock serializes sweeps across both roles. A successful API or worker leader holds it until that process stops, so no other replica can repeat a completed startup sweep over newly created runs. Followers retry acquiring leadership without a fixed attempt limit; if the leader exits or dies, either role can take over recovery. A failed or timed-out leader sweep releases the lock and retries up to three total sweep attempts per process. After the third failure, that process stops startup recovery and logs exhaustion; another replica can still acquire the lock. The advisory lock does not authorize mutations: durable run leases, coordinator plan claims, and child-dispatch reservations fence work across API and worker roles. Healthy child work on a surviving replica remains associated with its existing run identity. An expired coordinator child is restarted under that same child run ID after claiming its execution lease; a fresh lease held by another replica is skipped. When a coordinator terminates, its assemble-ready child and revision sandboxes are released after the final Scribe turn. A stopped coordinator releases them after the stop settles. On API startup and each coordinator heartbeat, a bounded page of terminal coordinators is revisited to recover claims left by interrupted cleanup. A child with a current durable execution lease or pending review is preserved; a terminal-parent child stuck Pending/InProgress without a live lease is reclaimed. A live preview retains its pod until preview expiry. Release reads one claim snapshot and uses Kubernetes UID/resourceVersion delete preconditions alongside the holder and lifecycle generation, so a replacement claim cannot be deleted by an older cleanup. Failures are logged as Terminal child cleanup warnings for retry rather than changing the coordinator's outcome. If a new revision stays pending on Kata, inspect these warnings and the child claim inventory before considering cluster capacity changes. An unbound AgentHost claim emits sandbox.provisioning_pending with the pod's PodScheduled=False reason when available; the coordinator displays that reason. If a pod cannot schedule before Sandbox:Kubernetes:AgentHostProvisioningTimeoutSeconds (default 600 seconds), the launch fails with the latest scheduling diagnosis and releases its claim. Check the pending pod's scheduler condition and node-pool autoscaler before retrying. Do not stop unrelated users' previews to free capacity. For a current preview, inspect GET /api/runs/{id} sandbox.current_binding: verified identifies the configured claim UID, Pod UID, namespace, generation, attempt, and source tree; unavailable or conflict includes a reason and must not be replaced with the historical sandbox.pod_name. A released execution lease and a retained child-owned preview are distinct lifecycle facts. A child's claim, Pod, and preview session do not attest its coordinator parent's claim or automatic preview. Older claims lacking the post-configure attestation remain explicitly unavailable. After a durable agent.turn.end, coordinator observation first waits Coordinator:PostTurnFinalizationGraceSeconds (default 10 seconds, clamped to 0.1–30 seconds) for assemble-ready or another terminal event. If the recovered child still owns an unexpired execution lease, it rechecks durable completion while that lease remains active, up to Coordinator:PostTurnFinalizationMaxWaitSeconds (default five minutes, clamped between the grace and ten minutes). An absent/expired lease or exhausted cap restores normal stall recovery; neither setting changes lease fencing. Inspect the child's execution lease and terminal run events before increasing the cap. Both API replicas can answer /api/ping without waiting for a sweep. A sweep has a five-minute deadline; followers and failed leader sweeps retry after 30 seconds, but a successful leader does not resweep. Look for Startup recovery sweep started, completed, exceeded, failed, or exhausted in API logs when diagnosing a restart. Readiness reflects workspace availability and successful initial static OAuth client reconciliation, not completion of the recovery backlog: operators should check the sweep log before assuming every interrupted run has been re-armed. /api/health, /healthz/workspace, and /oauth/* return 503 until the initial static OAuth client reconciliation succeeds; /api/ping stays responsive throughout. Failed reconciliations are logged and retried every five seconds after a 30-second attempt deadline without terminating the host. Database migrations and the bounded Copilot App registration validation still precede serving traffic.

AgentHost pre-delivery recovery diagnostics ​

Project agents, Assembly RAI, and Build & Test use warm-pool AgentHost claims. Before the first A2A request, the dispatch is generation-fenced and recovery is bounded:

  • agenthost_configure_copilot_token_refreshed means AgentHost explicitly rejected the configured Copilot credential, the API rotated that exact user/account scope, and Build & Test will recreate the one-time-configured pod once.
  • agenthost_configure_copilot_unauthorized means no different credential could be produced. The failure is not retried; the submitting user must repair GitHub/Copilot authorization.
  • Readiness, one-time configuration, a missing pod endpoint, or a reaped pod permits one fresh claim only while no model turn has been delivered.
  • Exhaustion writes one agent_host_unavailable terminal outcome with retryable: true. Retry creates a fresh run generation; the failed generation is not left live.

Logs include RunId, pod name, reason, recovery action, and bounded attempt counts. Token values are never logged, and the persisted terminal payload contains only the canonical error, retryability, and opaque dispatch id. Authorization/provider failures remain their specific typed errors rather than being converted to availability failures.

Once message:stream delivery starts, Agentweaver does not retry inside RemoteAgentProxy: request acceptance is uncertain, so replay could execute the same model turn twice. Post-acceptance transport and turn failures continue through the structured turn-failure reasons below. Operator Assistant conversations remain resumable; an AgentHost failure is recorded as a turn error and does not terminalize the conversation.

Diagnosing agent turn infrastructure failures ​

Agent turns executed through AgentHost use structured terminal reasons. agent_turn_internal_error with retryable: true is the fallback for an unstructured run.failed, an unsupported or unset A2A event, or a pod bridge turn that throws before emitting a structured terminal. Other A2A exceptions use a2a_transport_failure, whose retryability follows the transport error. A clean stream that ends without agent.turn.end uses retryable agent_host_turn_incomplete. In the collective Build & Test stage, the assembly reason prefixes the applicable reason with build_test_infra_, for example build_test_infra_agent_host_turn_incomplete.

This fallback does not hide more specific outcomes. Caller cancellation remains cancellation, and typed timeouts or failures retain their original error code and retryability. A retryable terminal means the workflow may safely consider a bounded retry or redispatch; it never converts the interrupted turn into a success.

For collective assembly, a typed retryable RAI provider or infrastructure failure receives one gate-only retry. The coordinator re-verifies the persisted aggregate tree hash and diff before retrying RAI; it does not rebuild integration, redispatch children, replace artifacts, or rerun completed Build & Test evidence. Content verdicts and revision feedback remain on the normal steering and human-review paths.

To investigate:

  1. Read GET /api/runs/{id}/terminal-diagnostic or run_failure_diagnostic.
  2. Start with observed_facts, then review supported_interpretations. Do not treat a nearby or repeated tool error as causal unless the terminal evidence directly references the same call or gate.
  3. Check evidence_sources and completeness. Missing telemetry does not erase durable terminal evidence, but partial or unavailable sources are not a healthy result.
  4. If denial_gate is present, repair the named authorization/configuration gate before retrying. A pending human approval is waiting, not denial.
  5. Follow the structured next_actions; they describe preconditions and effects but do not mutate the run, policy, or authorization.

Diagnostics in the run event are deliberately bounded, flattened to one line, and credential-redacted. They are safe context for triage, not a replacement for restricted server-side logs.

Cluster inventory collection uses the same explicit absence rule. Each inventory_sources entry reports available, no_resources, forbidden, timeout, unsupported, malformed, or collection_error. Only no_resources means collection completed successfully and found nothing. The Cluster page warns when any source is incomplete instead of presenting an unavailable inventory as an empty healthy one.

CommandPurpose
npm run azure:releaseFull semver release (see above)
npm run release:publishCreate the annotated tag and GitHub Release without deploying
npm run azure:deploy-from-release -- vX.Y.Z [--image-source acr-build]Deploy an existing published release (import already-published GHCR images by default, or rebuild from source)
npm run azure:provision-infraProvision/redeploy AKS, identity, monitoring, OAuth signing key, and PostgreSQL
npm run azure:deploy-from-localBuild, push, and verify images in ACR, then redeploy and cycle the warm pool
npm run azure:deploy-from-commit -- <sha-or-ref>Deploy an arbitrary exact commit through a temporary detached worktree
npm run azure:verifyVerify the current deployment

Use pnpm run in place of npm run if pnpm is your selected package runner. The runbook's individual-step section shows how to rerun one step.

Observability ​

Agentweaver ships with end-to-end telemetry using Azure Monitor OpenTelemetry Distro (Application Insights) and AKS Managed Prometheus.

Inspecting a transaction trace ​

Open a project, select Observability → Traces, then choose Preview trace for a coordinator run. The trace detail includes a timeline, span attributes, and persisted run events. Trace spans load in chronological pages; choose Load more spans until no more spans are available to inspect the complete trace. The opaque continuation keeps already loaded spans, selection, and tree state intact, and a failed page can be retried without reloading the whole trace. Tool spans carry bounded, redacted input/output previews in Application Insights, and persisted events are still loaded when the Events tab or a tool span needs additional context. It shows only trace data returned by Application Insights and the persisted run-event API. In particular, it shows the trace session ID only when the runtime emitted one, and it does not invent event timestamps when a legacy persisted event has no recorded time. The attributes pane is a fixed, safe schema rather than a dump of custom dimensions: it includes operational identity, model, provider, policy, sandbox, usage, status, and tool-payload capture state fields, but never prompts, credentials, raw tokens, secrets, unbounded output, or arbitrary tool payloads. For coordinator runs, child-run spans are grouped below the child agent that executed them using the persisted parent-run relationship; the original distributed trace parent remains available in the span data. See Transaction traces for the span and tool-call details.

If the Application Insights workspace is unavailable or slow, trace retrieval uses a 30-second server-side query budget by default, separately from the three-second dashboard-metrics budget. Set Metrics__AppInsights__TraceQueryTimeoutSeconds or APPINSIGHTS_TRACE_QUERY_TIMEOUT_SECONDS (1–60 seconds) to tune that trace budget. The trace panel automatically retries short-lived dependency failures with backoff, but surfaces a bounded query timeout without repeating the full long-running request. While a retry remains, the panel stays in its normal loading state instead of showing a failure banner. The API coalesces concurrent requests for the same run and cursor page into one bounded workspace query, and trace reads are not short-circuited by an unrelated dashboard-metrics cooldown. A recently retrieved page may be shown while the source recovers and is explicitly labeled as such; an unavailable source with no safe cached page is not presented as proof that the run has no trace data. Cursor paging remains incremental, so retry or Load more spans only requests the needed page. If the automatic attempts are exhausted, use Retry to start a fresh bounded trace load; platform operators can use the API log's query context and failure type to investigate workspace credentials, RBAC, and availability without logging KQL payloads.

Provisioning monitoring resources ​

Monitoring is provisioned as part of npm run azure:provision-infra. To rerun only that step, see the runbook's individual-step section (scripts/azure/steps/15-provision-monitoring.mjs).

This creates:

  • A Log Analytics workspace (agentweaver-logs)
  • A workspace-based Application Insights resource (agentweaver-insights) — workspace-based is required for the Agents (Preview) view
  • Stores the connection string as appinsights-connection-string in Key Vault
  • Enables AKS Managed Prometheus on the cluster

Finding the Application Insights resource ​

  1. Open the Azure Portal
  2. Navigate to your resource group (agentweaver-rg by default)
  3. Select the Application Insights resource named agentweaver-insights

Using the Agents (Preview) view ​

The Agents (Preview) view in Application Insights shows GenAI-specific telemetry including agent runs, token usage, and model calls.

  1. In the Application Insights resource, select Agents (Preview) from the left menu
  2. Use the time range picker to scope your investigation
  3. Filter by agent using the gen_ai.agent.name attribute — this maps to the configured agent name in the squad definition (e.g. morpheus, seraph)

Key span attributes emitted by Agentweaver:

AttributeDescription
gen_ai.agent.nameSquad agent name
gen_ai.agent.idAgent identifier
gen_ai.usage.input_tokensPrompt tokens consumed
gen_ai.usage.output_tokensCompletion tokens produced
gen_ai.request.modelModel deployment name
gen_ai.operation.namechat or execute_tool

To find all telemetry for a single run by its RunId:

  1. In Application Insights, select Search (or Transaction search)
  2. Enter the RunId (e.g. run_abc123) in the search box
  3. Alternatively, use Logs with a KQL query:
kusto
traces
| where customDimensions["RunId"] == "run_abc123"
| order by timestamp asc

Or to see all token usage for a run:

kusto
customMetrics
| where name == "agentweaver.token.usage"
| where customDimensions["run_id"] == "run_abc123"
| summarize totalTokens = sum(value) by tostring(customDimensions["agent_name"])

AKS Managed Prometheus metrics ​

Business metrics emitted by AgentWeaverMetrics are exported to the AKS Managed Prometheus workspace:

MetricTypeDescription
agentweaver_token_usage_totalCounterToken usage by agent and model
agentweaver_run_durationHistogramRun duration in milliseconds
agentweaver_run_errors_totalCounterRun errors by type
agentweaver_run_activeUpDownCounterCurrently active runs
agentweaver_run_queuedGaugeActive-project Ready backlog tasks awaiting coordinator pickup (backlog_tasks.state='ready' AND run_id IS NULL), sampled every 15s. Legacy name retained; aggregate with max, not sum, because every replica exports the same global snapshot.

To query in Azure Managed Grafana (linked to the Prometheus workspace), use standard PromQL:

promql
rate(agentweaver_token_usage_total[5m])

Worker autoscaling (queue depth vs. resource utilization) ​

k8s/base/worker-hpa.yaml scales agentweaver-worker between 2 and 3 replicas using CPU utilization at 70% and memory utilization at 80%. Neither metric is the Ready queue depth; resource utilization alone is an imperfect backlog proxy.

The agentweaver_run_queued gauge above (issue #108) exists specifically to provide a real queue-depth signal for this HPA. In the current system that signal is notruns.status='pending' — backlog pickup creates coordinator runs directly as in_progress. The durable queue is the set of active-project backlog tasks still in Ready with no bound run_id yet, which is what the gauge now publishes every 15 seconds. Because each replica exports the same shared-store total, Prometheus/KEDA queries must use max(agentweaver_run_queued) (or the equivalent single-series selector), not sum(...).

The HPA itself has not yet been switched over — wiring an external metric type into a plain HorizontalPodAutoscaler requires a Kubernetes External Metrics API adapter capable of serving Azure Monitor managed-Prometheus-backed queries, and no such adapter is currently provisioned in scripts/azure/. The two realistic paths forward (tracked against #108 — see decisions/inbox/niobe-108-hpa-investigation.md for the full analysis) are:

  1. KEDA with a Prometheus scaler (Microsoft's supported pattern for scaling on Azure Monitor managed Prometheus metrics) — query agentweaver_run_queued via the workspace's Prometheus query endpoint using max(...), not sum(...).
  2. Provision a k8s-prometheus-adapter-style External Metrics API adapter and wire worker-hpa.yaml with a type: External metric block pointing at it.

Until one of these is chosen and the supporting cluster component is provisioned, the worker continues to scale on CPU (with the gauge available for manual/Grafana-based capacity monitoring in the meantime).