Skip to content

Scaling Operations — Experience ​

This page is for the operator keeping Agentweaver moving as cluster load grows. The run and review contracts remain familiar, but replica restarts can interrupt connections and remote turns. The current checked-in deployment separates web and worker roles over Postgres; recovery uses leases and durable checkpoints, not a guarantee that every interrupted action is replayed invisibly.

For the reasoning behind these mechanics see the distributed execution & scaling deep dive; for the exhaustive schema, topology, and config details see the scaling data layer reference. Related operator context: Operations experience, Configuration, and the AKS deployment guide.

The mental model ​

The historical single-pod design combined API serving, orchestration, and a local SQLite database. That is context for the split, not the current Kubernetes deployment.

The current operator model separates:

  1. A managed database (Azure Database for PostgreSQL Flexible Server) replaces the SQLite file, so more than one process can write at once.
  2. Two pod roles replace the one combined pod: a web tier that talks to clients, and a worker tier that owns the actual runs.
  3. A lease lets many worker pods share the pool of runs without ever stepping on each other.
  4. Durable event polling keeps a run watchable from any web pod, even when a different worker pod is executing it.

The operator's job after scaling is mostly about the second and third points: making sure enough web pods exist for request load, enough worker pods exist for run backlog, and that runs are being leased and renewed cleanly.

What scaling looks like in practice ​

More pods, in two roles ​

Instead of one Deployment you operate two, both built from the same image and told apart by a role flag:

  • Web pods serve REST, authentication, and live event streams without owning durable run execution. The manifest starts with two replicas. Restarts can disconnect clients; persisted events support reconnect and catch-up.
  • Worker pods claim runs, drive orchestration, dispatch sandbox leaf execution, and write durable checkpoints/events. The active HPA has minimum 2, maximum 3, with CPU 70% and memory 80% utilization targets. Backlog pressure is useful operational context, not the active HPA input.

k8s/base/worker-hpa.yaml contains a commented future KEDA alternative and a commented web HPA example; neither is deployed by that file. The exported agentweaver_run_queued gauge counts eligible Ready backlog tasks in active projects awaiting pickup. It is not the count of all leased or running runs. Keep workers available; scaling them to zero does not make web-role pods take over execution.

A managed database instead of a file ​

The SQLite file and its single-writer disk are gone. In their place is a managed Postgres Flexible Server reached privately from inside the cluster, with zone-redundant high availability and managed point-in-time backups. Two practical consequences for you:

  • The old /data ReadWriteOnce disk and the SQLite backup CronJob are retired — backups are now the database's managed responsibility.
  • The /workspace shared volume stays. It is multi-attach-safe and still holds the git worktrees that worker and sandbox pods share, so it is not a scaling bottleneck.

The database connection is passwordless: pods authenticate using the cluster's workload identity, so there is no DB password to store or rotate in a secret.

What stays consistent for end users ​

The run/review model and REST/event-stream contracts do not change with replica count. This is contract continuity, not a promise that users cannot notice reconnections, delayed scheduling, or failed turns.

The current Overview uses Recent projects, AI usage & performance, Activity feed, and Needs attention. Those projections help users find work; they do not demonstrate lease or fencing guarantees.

Live watching crosses replicas. A run executes on a worker while the browser may connect to another web replica. EfRunEventStream appends durably and subscribers poll the shared event table from their cursor (a 250 ms polling interval). This is not a deployed LISTEN/NOTIFY relay. Connection loss can delay delivery; reconnect/catch-up is distinct from replaying an interrupted model turn.

How runs survive replica restarts ​

This is the behavior that most changes the operator's day, and it is worth understanding well.

An active workflow watcher claims a lease before processing its stream. Leasing lets workers coordinate ownership; restart recovery separately decides what can resume. A lease records its owner, expiry, fencing token, and heartbeat.

Claiming is atomic. A guarded database update succeeds only if a run is free or its lease has expired. Competing claims have one winner; each successful acquisition advances the fencing token. This protects ownership, not arbitrary external side effects from an interrupted turn.

Now the restart story:

  • A worker pod is restarted, drained, or crashes.
  • It stops renewing the heartbeats on the runs it held, so those leases lapse.
  • A subsequent eligible watcher can claim a free or expired lease with the guarded update.
  • Recovery depends on the run state: AwaitingReview can resume from a usable checkpoint; interrupted coordinator parents recover through their persisted work plan. Stranded in-progress child turns become retryable a2a_transport_interrupted failures for coordinator redispatch, while stranded root turns fail as stranded_in_progress. Lease expiry alone does not replay a model turn.

A graceful shutdown stops new claims and attempts to release owned leases. The worker disruption budget keeps at least one worker available; it is not a guarantee that every in-flight turn finishes before termination.

Renewal and release match both owner and fencing token. Terminal handlers check current lease ownership before updating run state, rejecting a stale owner's terminal outcome. Do not broaden those guards into a claim that every filesystem, tool, or external side effect is fenced.

The operator's mental checklist ​

When you operate a scaled Agentweaver, these are the things worth watching:

Open project Diagnostics → Global for health checks and counts, then Cluster for claims and resource readiness. Interpret process uptime separately from fleet-wide persisted run counts.

  1. Web replicas vs request load. Inspect request/connection pressure before changing the two-replica baseline; do not assume a web HPA is installed.
  2. Worker HPA and backlog. Check CPU/memory targets and the 2–3 replica bounds alongside eligible Ready backlog depth. A growing backlog can also indicate unavailable projects or scheduling delays.
  3. Leases are being renewed. Healthy workers refresh heartbeats; runs whose leases keep expiring and getting re-claimed point at workers that are crashing, starved, or being killed too aggressively.
  4. Database health. Postgres is the shared source of truth. Watch availability, connection headroom, and event-read latency. The current cursor-polling relay does not depend on session-bound LISTEN/NOTIFY connections.
  5. Roll one worker at a time. Worker disruption budgets and graceful drain are tuned so leases hand off cleanly; respect them during upgrades so in-flight runs checkpoint and migrate rather than restart.

Historical rollout and rollback boundaries ​

The P1/P2/P3 plan explains how the architecture developed; it is not a pending rollout or a rollback runbook:

  • P1 — remote leaf execution: moves heavyweight sessions into sandbox pods. in-api selects local execution when deployed, but does not migrate active remote turns.
  • P2 — shared Postgres: removes the single-writer deployment constraint. Switching Database:Provider does not copy or reconcile data; rollback requires an explicit data restoration or migration plan plus compatible storage and replica settings.
  • P3 — web/worker roles and leases: separates serving from execution. Web-role pods do not automatically become workers when worker replicas reach zero; a role/topology rollback must keep an execution owner available.

Apply reviewed configuration and manifests through the normal deployment process. Validate ownership, checkpoint recovery, storage compatibility, and capacity before calling any rollback safe.

Diagram details and constraints
ElementContract
titleScale roles, keep state shared
takeawayWeb serves requests; workers execute; Postgres coordinates durable state; pods run leaves.
group-title-0CONTROL-PLANE ROLES
group-title-1SHARED STATE AND LEAF COMPUTE
BrowserBrowser
BrowserREST and watch client
Browserrequest / SSE response
BrowserConnects to web replicas; not a worker-local queue.
Web tierWeb tier
Web tierRequests and event reads
Web tierbase: 2 replicas
Web tierUses shared durable state; reads event cursors.
Worker tierWorker tier
Worker tierExecution and ownership
Worker tierHPA: 2-3 replicas
Worker tierCPU 70% / memory 80%; leases and workflow state.
Current boundaryCurrent boundary
Current boundaryConfiguration, not a probe
Current boundaryKEDA / web HPA: examples
Current boundaryChecked-in scaling settings are not live cluster evidence.
PostgresPostgres
PostgresShared durable state
Postgresevents / leases / checkpoints
PostgresEvent relay polls the table; not LISTEN/NOTIFY.
Sandbox podsSandbox pods
Sandbox podsRemote AgentHost leaves
Sandbox podsrun context + A2A
Sandbox podsReturn events to the worker; no direct pod DB access.
e0requests
e1state / events
e2persist
e3execute
e4results
noteWorker HPA is active; KEDA and web HPA blocks are commented proposals. Kubernetes owns scheduling.
n0Connects to web replicas; not a worker-local queue.
n1Uses shared durable state; reads event cursors.
n2CPU 70% / memory 80%; leases and workflow state.
n3Checked-in scaling settings are not live cluster evidence.
n4Event relay polls the table; not LISTEN/NOTIFY.
n5Return events to the worker; no direct pod DB access.
groupsCONTROL-PLANE ROLES; SHARED STATE AND LEAF COMPUTE
Diagram details and constraints
ElementContract
titleLease transfer is not turn replay
takeawayAn expired owner can be replaced; continuation depends on the persisted run state.
group-title-0LEASE OWNERSHIP
group-title-1STATE-DEPENDENT RECOVERY
Worker AWorker A
Worker AGuarded lease claim
Worker Aowner A + token n
Worker ARenewals must match both owner and fencing token.
Lease expiresLease expires
Lease expiresRenewals cease
Lease expiresfree / expired eligibility
Lease expiresFailure is not a message; expiry permits a later claim.
Worker BWorker B
Worker BWin eligible next claim
Worker Bowner B + token n+1
Worker BOld-token terminal ownership checks reject stale results.
Visible failureVisible failure
Visible failureStranded child / root
Visible failureretryable child vs root
Visible failureChild may be redispatched; no automatic mid-turn replay.
Inspect run stateInspect run state
Inspect run stateCheckpoint or work plan
Inspect run staterecovery prerequisites
Inspect run stateAwaitingReview, coordinator, and stranded turns differ.
Durable recoveryDurable recovery
Durable recoveryEligible saved state
Durable recoverycheckpoint / coordinator plan
Durable recoveryRecover only supported state; missing checkpoints surface.
e0renewals stop
e1next claim
e2inspect
e3stranded turn
e4recoverable
noteFencing protects guarded ownership operations, not every external tool side effect.
n0Renewals must match both owner and fencing token.
n1Failure is not a message; expiry permits a later claim.
n2Old-token terminal ownership checks reject stale results.
n3Child may be redispatched; no automatic mid-turn replay.
n4AwaitingReview, coordinator, and stranded turns differ.
n5Recover only supported state; missing checkpoints surface.
groupsLEASE OWNERSHIP; STATE-DEPENDENT RECOVERY