Skip to content

Testing Strategy — Conceptual Deep Dive ​

Purpose and mental model ​

Agentweaver is verified as a control system, not as a collection of isolated helpers. The core risk is not just "does this method return the right value?" It is "can an AI-driven run move through a controlled lifecycle without escaping its workspace, skipping review, losing events, corrupting repository state, or trusting the wrong identity?"

The test strategy therefore mirrors the architecture from the system overview:

  1. Intent plane tests check HTTP and MCP-facing contracts: authentication, project APIs, review decisions, coordinator gates, memory tools, and staging smoke paths.
  2. Control plane tests check durable state and orchestration: run stores, event streams, workflow gates, coordinator planning, merge coordination, recovery, and decision ledgers.
  3. Execution plane tests check governed side effects: filesystem containment, shell sandbox policy, git worktrees, provider seams, and agent tool behavior.

The philosophy is conservative: keep the default suite fast and hermetic, but exercise real infrastructure boundaries wherever correctness depends on them. Tests replace live model calls and external GitHub/network dependencies with deterministic seams, while still using real HTTP routing, real SQLite databases, real git repositories, real workflow state machines, and real sandbox path logic.

Where this lives: tests/Agentweaver.Tests, tests/Agentweaver.Tests/Helpers, tests/e2e, docs/deep-dive/00-system-overview.md.

The testing pyramid ​

Agentweaver's pyramid is intentionally wide in the middle. Pure unit tests are useful for algorithms and validators, but many of the important failures happen at service boundaries: middleware ordering, database state transitions, workflow gates, stream persistence, git branch updates, and sandbox policy wiring. Those are covered with in-process integration tests rather than brittle end-to-end tests.

LayerWhat it provesTypical dependencies
Unit/componentLocal rules are deterministic: path validation, PKCE math, token validation, command construction, parsers, stores, DAG frontier logic.In-memory objects, fake HTTP handlers, isolated temp directories.
Store/infrastructureDurable contracts hold: SQLite schemas, compare-and-swap transitions, append-only records, replay, idempotent persistence.Real SQLite files or in-memory SQLite connections.
API integrationThe real host routes requests through middleware, auth, endpoint mapping, services, stores, and response serialization.WebApplicationFactory<Program>, temp SQLite, temp worktree/checkpoint roots, test auth keys.
Workflow/service integrationIndividual graph, review, merge, and recovery contracts without live model calls; coverage depends on the fixture and invocation.Fake workflow agents, real git/stores where relevant, pending gates and async polling.
Security/regressionKnown escape and race classes stay closed.Real filesystem paths, real git merges, sandbox validators, opt-in live-provider canaries.
FrontendBrowser state, reducers, routing and components under deterministic inputs.Vitest/Testing Library in apps/web.
PostgreSQL integrationReal provider transactions, sequence serialization and migrations.PostgreSQL Testcontainers; not proved by SQLite-backed EF tests.
MCP real-processActual MCP transport/validation against controlled issuer/API inputs.Loopback MCP process, synthetic JWKS and stub API; not a deployed end-to-end environment.
E2E smokeSelected deployed-user paths.Playwright against configured AKS_BASE_URL or staging; credentials/environment required.

The shape is a trade-off. The suite avoids depending on live providers by default because model output is nondeterministic and credentials are sensitive. But it also avoids over-mocking the system: most integration tests boot the real API and then replace only the external seams that would make the test slow, flaky, or non-hermetic.

How API integration tests host the system ​

The main integration pattern is a custom WebApplicationFactory<Program>. Each factory starts the real ASP.NET Core application in-process, then injects test configuration before startup:

  • a unique SQLite database path;
  • isolated worktree and checkpoint directories;
  • test bearer keys and users;
  • test provider configuration values required at startup;
  • explicit test authentication/bypass seams when identity is not the subject;
  • service replacements for live external seams.

This approach is important because middleware order and DI wiring are part of the contract. A project endpoint test is not merely testing a service method; it verifies that authentication, route binding, JSON naming, service registration, persistence, and response status codes all agree.

There are several specialized factories because different subsystems need different seams:

  • General API factory: isolated storage/roots, dummy provider values and an explicit auth test seam.
  • Projects factory: real filesystem provider with a no-op Git initializer for CRUD-focused tests.
  • Review factories: controlled identities and repositories for authorization and branch effects.
  • Workflow factory: deterministic runner and workflow-agent replacements; the calling test determines which graph path executes.
  • Coordinator factory: deterministic planning/classification and disabled auto-dispatch for scoped phases.
  • OAuth/MCP fixtures: controlled issuer, signing/token state and HTTP/process peers appropriate to each test.

Historical fixture configuration keys are test setup, not evidence of production API-key or ambient GitHub-token compatibility.

The common principle is: replace the world outside Agentweaver, not Agentweaver's own control plane.

Where this lives: tests/Agentweaver.Tests/Helpers.

Fakes, fixtures, and real dependencies ​

Agentweaver tests use fakes deliberately, not casually. A fake is acceptable when it stands at a nondeterministic or external boundary and preserves the shape of the production contract. A fake is not used to skip the behavior being tested.

Key patterns:

  • SQLite is real. Many tests use unique database files, and some EF-backed coordinator or memory tests use an in-memory SQLite connection. This catches SQL schema, transaction, uniqueness, append-only, and compare-and-swap behavior that a mock would miss.
  • Git is real when git behavior matters. Review, commit, merge, integration-branch, and workflow tests create temporary LibGit2Sharp repositories. They assert branch tips, tree hashes, merge conflicts, and worktree cleanup instead of assuming a git abstraction worked.
  • Agents are deterministic. TestFileEditAgentRunner can make a real file change, make no change, or simulate content-safety failure. FakeWorkflowAgentFactory keeps RAI, Rubberduck, and Scribe turns deterministic so tests can focus on workflow routing.
  • Coordinator drafting is deterministic. The coordinator suite replaces model drafting with a fake drafter that produces stable outcome specs from the goal and memory context. The subsequent gates, persistence, and orchestration are still real.
  • Network calls are stubbed at the HTTP boundary. OAuth, GitHub token, org membership, and API-tool tests use fake handlers or in-memory stores so tests can exercise success, denial, SAML, rate-limit, and error semantics without live GitHub.
  • Sandbox tests use real filesystem objects. Path tests create actual directories, files, and, when permitted by the OS, symlinks. Escape tests rely on canary files outside the sandbox as the oracle.

This split keeps the suite reliable while still testing the failure modes that historically matter: database races, branch divergence, path traversal, middleware auth behavior, and workflow gate consumption.

What each subsystem's tests are trying to protect ​

Memory and decisions ​

Memory tests treat the memory system as a ledger. They verify that inbox submissions are idempotent by slug, merge creates a canonical decision and marks the source entry, rejection retains audit state without creating a decision, filters return the expected status/type/agent slices, and exported/imported data stays tied to the project model.

The tests also cover the agent-facing API tools. Those tools must return useful strings for recoverable conditions such as conflicts or server errors instead of throwing opaque tool-execution failures. That matters because an agent turn should be able to learn "this decision was already recorded" and continue.

Rebuilder rule: test both the HTTP ledger and the agent tool facade. A memory system is only useful if humans can inspect it and agents can write to it without destabilizing runs.

OAuth, auth, and MCP ​

Auth tests are split between pure protocol rules and hosted middleware behavior:

  • token signing produces audience-bound JWTs and rejects tampering or wrong audiences;
  • redirect URI policy allows loopback/native-client shapes and rejects unsafe destinations;
  • PKCE requires S256 and rejects missing or weak challenge inputs;
  • authorization codes and refresh-token paths are specified as single-use/rotating where implemented;
  • Entra and endpoint-classified policies reject invalid identities or unauthorized project roles;
  • MCP accepts broker tokens and rejects upstream tokens/static API keys;
  • discovery and metadata routes are checked through in-process or staging smoke paths.

Executable coverage includes AuthenticationSchemeCutoverTests, OpenIddictAuthorizationServerTests, and McpBrokerRealProcessTests. A skipped acceptance scenario is not evidence that its behavior passed, and old fixture keys are not supported production authentication modes.

Rebuilder rule: test OAuth as a state machine, not just as JSON metadata. Codes, verifiers, redirect URIs, aud claims, refresh rotation, and revocation are security invariants, so each should have a positive and negative test.

Projects and GitHub integration ​

Project tests cover CRUD, deletion, workspace selection, role boundaries, capability redaction, and caller-bound GitHub repository-selection codes. Direct repository input is rejected by GitHubRepositorySelectionEndpointsTests; a GitHub connection is not platform sign-in.

Most project endpoint tests do not perform real clones; they use a no-op initializer and isolated workspaces. Git behavior is reserved for tests where branch state is the point.

Rebuilder rule: project tests should distinguish "metadata and policy" from "actual git mutation." Use a fake initializer for project CRUD, but use real repos for merge, worktree, and branch invariants.

Single-run workflows, review, and merge ​

Graph, service, and review tests cover parts of the central lifecycle:

  1. a run starts against a repository and branch;
  2. the agent changes an isolated worktree;
  3. the workflow reaches review with a diff and tree hash;
  4. a valid reviewer approves, declines, or requests changes;
  5. merges serialize and update the originating branch only when allowed;
  6. terminal states clean up or preserve worktrees according to outcome.

Tests assert observable state, pending gates, tree hashes, cleanup, diffs, events, and review consumption where those are in scope. In particular, current WorkflowIntegrationTests checks that retired public POST /api/runs is rejected (401 or 410); its name does not prove an agent-to-review-to-merge traversal.

The important race tests use compare-and-swap style assertions. For example, concurrent approve and request-changes attempts must result in exactly one winner. Append-only revision tests make direct database tampering fail. Prompt-injection tests verify reviewer feedback is nonce-fenced before it is handed back to the agent.

Rebuilder rule: never test review as a boolean flag only. Test the whole boundary: owner identity, pending gate, state transition, branch mutation or non-mutation, event emission, idempotency, and cleanup.

Coordinator flows ​

Coordinator tests are heavy because coordinator correctness is mostly state-machine correctness. They cover:

  • outcome-spec draft, persistence, event emission, and suspension at a confirmation gate;
  • confirm, revise, decline, owner-scoping, missing-run, and no-pending-gate outcomes;
  • bounded waiting for a gate that arms shortly after the UI observes awaiting_confirmation;
  • deterministic work-plan persistence after confirmation;
  • subtask dependency frontier rules;
  • child observation, failure routing, retry, pickup ownership, roster dispatch filters, and assembly finalization;
  • integration-branch construction from child outputs, including conflict handling;
  • coordinator event persistence and replay contracts.

The coordinator suite uses deterministic planning seams because a model-generated plan would make the suite nondeterministic. The system behavior after the plan exists is still tested through real stores and services.

Rebuilder rule: test the coordinator as a durable DAG plus gates. A good suite should be able to restart, recompute the frontier, observe child states, assemble only eligible outputs, and reject double decisions.

Sandbox and tool governance ​

Sandbox tests are intentionally adversarial. They check path traversal, absolute paths, Windows drive paths, null bytes, symlinks, search traversal patterns, excluded directories, governance default-deny behavior, command validation, shell executor command construction, and regressions where a provider once had its own unsafe resolver.

The live provider sandbox-escape tests are opt-in. They create an in-sandbox marker and an out-of-sandbox canary, instruct the provider-backed agent to read both, and pass only if:

  • the in-sandbox marker can be read;
  • at least one out-of-sandbox attempt is denied;
  • the canary never appears in the final response, streamed events, or logs;
  • tool call/result/error events are correlated by call id.

That test is expensive and credentialed, so it is not part of the default suite. But it is the right shape for proving the provider path is wired through governance rather than only testing local helpers.

Rebuilder rule: sandbox tests need a positive control and a negative oracle. A test that only asserts "denied" can be vacuous; a test that proves safe access works and unsafe canary access does not leak is stronger.

Events, persistence, and recovery ​

Event tests protect the "durable event log is truth" invariant:

  • appending writes through to SQLite before returning;
  • a new stream instance can replay events after simulated restart;
  • cursor resume returns only events after the requested sequence;
  • persisted coordinator streams stay ordered and idempotent;
  • terminal events survive long enough for observers and replay consumers.

Recovery and watch-loop tests build on that by checking terminal output, restart services, pending gates, and run status reconciliation.

Rebuilder rule: do not treat streams as only live websockets. Test the database rows directly, then test replay through the public stream abstraction.

E2E smoke ​

Playwright targets a configured deployed instance; it does not automatically launch a fake local stack. Evaluate each test's enabled/skipped status and credential requirements. Entra sign-in and broker-token MCP are the current contracts; old GitHub-login/static-key scenarios are not compatibility promises. Executable loopback OAuth/MCP tests remain a separate assurance boundary.

Rebuilder rule: keep E2E small. Use it to prove deployment wiring and the most important user-visible paths, not to duplicate every API integration test.

Determinism in agent and coordinator tests ​

Agentweaver cannot make a model deterministic, so tests put determinism at the seam immediately outside the model. The fake agent still performs real file operations and emits representative events; it just chooses from known modes.

Deterministic seam inputWhat the fake suppliesWhat a consuming test must assert separately
File-edit modeA real fixture file changeDiff/tree identity and whichever review/merge path the test invokes.
No-change modeAn unchanged resultThe configured graph's no-change outcome, not a universal production state transition.
Safety-failure modeA synthetic content-safety exceptionFailure classification and cleanup in the tested service/graph.
Review decisionAn explicit fixture responseAuthorization, pending-gate correlation, CAS consumption, and branch effects.

This pattern has two advantages:

  1. Tests can retain real graph/store/Git behavior while replacing model output; the fixture determines which of those boundaries actually runs.
  2. The test has a stable oracle. If a run fails to reach awaiting_review, there is a control-plane problem, not a model-quality problem.

Coordinator tests use the same idea. The drafter is deterministic, but the persisted spec, gate consumption, work-plan records, dependency edges, and events are real. This is the right compromise for a system where the model is one participant, not the source of authority.

In-memory versus real dependencies ​

The suite uses a simple decision rule:

Examples:

  • SQLite is real because transactionality, uniqueness, append-only behavior, and compare-and-swap transitions are test subjects.
  • Git repositories are real when branch and merge behavior are test subjects.
  • GitHub and model providers are fake by default because live network and model output are not the test subject in most runs.
  • The API host is real because route, middleware, DI, and serialization wiring are part of the contract.
  • The sandbox path resolver is real because path containment is the contract.

This rule is what a rebuild should preserve. The exact class names can change; the boundary logic should not.

Deliberate test boundaries ​

The suite is strong around control-plane invariants, but there are deliberate gaps:

  • Live model-provider behavior is not part of the default suite. Provider-backed sandbox escape tests exist, but they are opt-in through environment configuration.
  • Staging Playwright tests are smoke tests, not full workflow coverage.
  • Skipped acceptance cases are not passing evidence; executable OAuth lifecycle and real-process MCP tests cover distinct controlled boundaries.
  • Kubernetes sandbox execution is not proven by a default live-cluster E2E. The default coverage focuses on command construction, policy behavior, and API-side sandbox contracts.
  • Frontend unit/component tests live separately in apps/web; the .NET test directory is not their coverage boundary.
  • Performance, load, and long-running multi-agent soak behavior are not represented as a normal test layer.

The documented scope includes .NET, frontend Vitest, PostgreSQL integration, controlled MCP process tests, and deployed Playwright smoke. None by itself proves live model, Kubernetes, or multi-agent soak behavior.

These gaps are acceptable only if they are explicit. The default suite should remain hermetic, but release gates should add opt-in live checks for the boundaries that cannot be proven locally.

Invariants a rebuild should preserve ​

A rebuilt Agentweaver should have tests that protect these invariants:

  • Authority stays in the control plane. Clients and MCP tools request actions; they do not directly mutate run state, memory ledgers, or branches.
  • Runs are durable and replayable. Every run has a durable identity, monotonic event sequence, and recoverable terminal story.
  • Work happens in isolated worktrees. The protected branch is unchanged until review/merge logic explicitly advances it.
  • Review decisions are scoped and consumed once. Owner checks, pending gates, idempotency, and compare-and-swap transitions prevent stale or double decisions.
  • Merge failures are safe. Conflicts do not advance the originating branch, and conflict details do not leak raw file content in unsafe places.
  • Sandbox boundaries fail closed. Unknown tools, suspicious paths, weak executors, symlink escapes, and out-of-root operations are denied before side effects.
  • Safe operations still work. Tests must prove agents can read/write/search inside the workspace so denial checks are not vacuous.
  • Auth is explicit. Entra/broker credentials and endpoint/project roles are checked, OAuth redirects are constrained, S256 PKCE is required, and production bypasses are guarded.
  • Memory promotion is atomic. Inbox entries, decisions, session context, and exported files remain consistent enough for future agents to trust.
  • Coordinator work is a DAG. Dependencies, child states, retries, assembly, and collective review are durable and recomputable.
  • External nondeterminism is isolated. Tests can run without live models, real GitHub, or staging unless the test is explicitly opt-in.

How to structure tests when rebuilding ​

Start with the contracts, then choose the lightest dependency that can prove each one.

  1. Write pure tests for pure rules. Path normalization, redirect validation, PKCE challenge matching, token validation parameters, command builders, parsers, and DAG frontier logic should be fast and table-driven.
  2. Use real SQLite for persistence semantics. If a feature depends on transactions, unique indexes, append-only triggers, replay, or compare-and-swap, do not mock the store.
  3. Use in-process API tests for route contracts. Boot the real host with temp configuration and make HTTP requests through HttpClient.
  4. Use real git for branch claims. If a test says "branch unchanged" or "merge conflict," assert against actual git objects.
  5. Replace only external seams. Fake model turns, GitHub HTTP responses, credential stores, and project initializers, but keep Agentweaver's services real.
  6. Poll asynchronous workflows through observable state. Do not sleep blindly. Poll run status, pending gates, stream events, or database rows with bounded timeouts.
  7. Make security tests non-vacuous. Include a safe positive control and a denied negative control, preferably with a canary value that must not leak.
  8. Keep E2E narrow and honest. Use it for deployment wiring, browser redirects, and staging smoke; keep detailed behavior in hermetic integration tests.

The rebuild target is not identical file names. It is the same confidence model: deterministic tests around nondeterministic agents, real persistence for durable claims, real git for repository claims, adversarial tests for security boundaries, and small opt-in live checks for everything that cannot be proven offline.

Diagram details and constraints
ElementContract
titleTesting boundary · choose what must stay real
takeawayReplace nondeterministic dependencies without replacing the behavior under test.
group-0-titleREAL BEHAVIOR / INFRASTRUCTURE
group-1-titleCONTROLLED EXTERNAL SEAMS
Real API hostReal API host
Real API hostWebApplicationFactory / Program
Real API hostMiddleware, DI, binding and response contracts
Real API hostAgentweaverWebApplicationFactory
Deterministic agent seamDeterministic agent seam
Deterministic agent seamTestFileEditAgentRunner
Deterministic agent seamReal file write / no-change / synthetic safety error
Deterministic agent seamTestFileEditAgentRunner:44–80
Real SQLite storesReal SQLite stores
Real SQLite storesRaw and EF-backed test databases
Real SQLite storesTransactions, uniqueness, state and CAS behavior
Real SQLite storesProjectsWebApplicationFactory
Controlled HTTP boundaryControlled HTTP boundary
Controlled HTTP boundaryStub handlers / issuer / API
Controlled HTTP boundaryExternal dependency shape, not live upstream proof
Controlled HTTP boundaryMcpBrokerRealProcessTests
Real Git / policy logicReal Git / policy logic
Real Git / policy logicRepositories and worktree operations
Real Git / policy logicBranches, tree hashes, containment and validators
Real Git / policy logicWorkflowWebApplicationFactory
Planning / workflow seamsPlanning / workflow seams
Planning / workflow seamsFixture-specific factory replacements
Planning / workflow seamsSome fixtures suppress dispatch or Rai/Scribe
Planning / workflow seamsCoordinatorWebApplicationFactory
Real PostgreSQL fixtureReal PostgreSQL fixture
Real PostgreSQL fixturepostgres:16-alpine Testcontainer
Real PostgreSQL fixtureProvider locks/migrations need the actual engine
Real PostgreSQL fixturePostgresFixture:9–29
Real MCP processReal MCP process
Real MCP processLoopback executable test
Real MCP processSynthetic JWKS + stub API remain controlled
Real MCP processMcpBrokerRealProcessTests:235–255
Real API hostinject seam
Real API hostplanning
scopeScope varies by fixture. These tests do not establish live Entra, model-provider or Kubernetes behavior.
groupsREAL BEHAVIOR / INFRASTRUCTURE; CONTROLLED EXTERNAL SEAMS
Diagram details and constraints
ElementContract
titleTest coverage · independent layers, distinct proof
takeawayEach layer proves a different seam; this is a coverage map, not a sequential execution pipeline.
group-0-titleDETERMINISTIC / IN-PROCESS
group-1-titlePROVIDER / PROCESS / AUTH / LIVE
Unit / componentUnit / component
Unit / componentDomain services and validators
Unit / componentFocused logic without live provider dependencies
Unit / componenttests/Agentweaver.Tests
Frontend VitestFrontend Vitest
Frontend VitestHooks, reducers and components
Frontend VitestSeparate web suite; not downstream of .NET tests
Frontend Vitestapps/web/package.json:11
In-process APIIn-process API
In-process APIFactory + real Program host
In-process APIRouting / identity policy / JSON / persistence
In-process APIHelpers/*WebApplicationFactory
Workflow / Git / policyWorkflow / Git / policy
Workflow / Git / policyDeterministic execution seams
Workflow / Git / policyReal Git when needed; do not infer live model proof
Workflow / Git / policyTestFileEditAgentRunner:44–80
PostgreSQL integrationPostgreSQL integration
PostgreSQL integrationTestcontainers + migrations
PostgreSQL integrationProvider-specific locks and concurrency behavior
PostgreSQL integrationPostgresFixture:9–29
Real MCP processReal MCP process
Real MCP processLoopback subprocess tests
Real MCP processControlled issuer/JWKS and stub API route
Real MCP processMcpBrokerRealProcessTests
OAuth server testsOAuth server tests
OAuth server testsConsent, PKCE and refresh state
OAuth server testsActive focused tests; not blanket skipped OAuth
OAuth server testsOpenIddictAuthorizationServerTests
Opt-in / staging browserOpt-in / staging browser
Opt-in / staging browserPlaywright deployment tests
Opt-in / staging browserLive target + explicit environment prerequisites
Opt-in / staging browsertests/e2e/playwright.config.ts
prerequisitesPrerequisites differ: .NET / Node locally; Docker for Postgres; a configured live target for staging.
scopeNo arrows: layers run independently. Source presence is evidence of coverage, not a passing test run.
groupsDETERMINISTIC / IN-PROCESS; REAL BOUNDARIES / OPT-IN
Diagram details and constraints
ElementContract
titleAPI test hosting · configure, substitute, request
takeawaySpecialized factories keep Program wiring real while selecting isolated state and controlled seams.
group-0-titleFACTORY SETUP
group-1-titleHOST / REQUEST / ASSERT
Test caseTest case
Test caseSelect a specialized factory
Test caseChoose the subsystem and behavior to preserve
Test caseHelpers/*WebApplicationFactory
Fixture configurationFixture configuration
Fixture configurationIsolated DB, worktree, checkpoints
Fixture configurationRequired provider settings; explicit identity seams
Fixture configurationWorkflowWebApplicationFactory:34–60
Service replacementService replacement
Service replacementSwap IAgentRunner / agent factory
Service replacementDeterministic writes; some gates short-circuited
Service replacementWorkflowWebApplicationFactory:61–82
Real Program hostReal Program host
Real Program hostStartup + DI + middleware
Real Program hostNot a direct service-method-only test
Real Program hostWebApplicationFactory
Isolated real stateIsolated real state
Isolated real stateSQLite + filesystem / Git as needed
Isolated real stateDispose fixture-owned databases and directories
Isolated real stateWorkflowWebApplicationFactory:86–111
Test HTTP clientTest HTTP client
Test HTTP clientRoute, JSON and identity contract
Test HTTP clientRequest traverses the configured host
Test HTTP clientWorkflowIntegrationTests:20–31
Controlled executionControlled execution
Controlled executionNo live model invocation required
Controlled executionFake preserves relevant input/output shape
Controlled executionTestFileEditAgentRunner:44–80
Assertions + cleanupAssertions + cleanup
Assertions + cleanupStatus / payload / state as applicable
Assertions + cleanupRetired POST /api/runs test accepts 401 or 410
Test caseconfigure
Test casereplace
Fixture configurationbuild
Service replacementinject
Controlled executionwrites
Test HTTP clientassert
scopeFixture legacy auth settings are not production architecture. The retired-route test proves no full workflow.
groupsFACTORY SETUP; HOST / REQUEST / ASSERT