Autonomous QA for Claude Code

AI writes code in seconds.
Catching what it broke shouldn't take all afternoon.

A background bridge that captures the intent behind your prompt, runs headless browser journeys on every change, and produces video proof + a ready-to-paste fix before you merge.

Install GitHub App
Zero manual test writing
Encrypted test accounts
SHA-based freshness cache
closed-loop-qa.sh
Bridge Connected (daemon → platform)
Coding Agent (Claude Code / Cursor)
branch: feat/quick-checkout

> user: "Add Apple Pay button to checkout form"

> agent: Modifying components/pay-button.tsx...

[Bridge Intent Stream Captured]
• Intent: Bypass address form on Apple Pay token
• Rationale: Accelerate checkout completion
• SHA: f1e2d3c4 (Pushed to platform)
$ agent-bridge check
✓ Verification enqueued. Booting sandbox + headless browser...
Autonomous QA Sandbox
High Risk Flagged
Analysis: Touched checkout flowRisk: High
Seeded Auth Journey: /login -> passed (190ms)
Nested Scroll: Dashboard container scrolled (element-targeted)
Journey 'checkout' FAILED
Invoice creation error: customer_address missing
video.webm (session recording)trace.zip (DOM snapshots)
Forensic Proof Produced: Copy-paste this remediation prompt back to Claude Code to fix in seconds.

Why traditional CI fails AI coding workflows

When developers generate features at 10x speed, testing blindspots compound instantly.

01 / ADJACENT REGRESSIONS

Nobody re-tests unmentioned flows

Your coding agent fixes an input validation rule, but silently breaks a checkout invoice flow 3 screens downstream that nobody wrote a test for.

02 / REPRODUCTION WASTE

Hours spent reproducing stack traces

CI reports a red failure with a cryptic stack trace — no video replay, no network waterfall, and zero record of why the code was changed.

03 / MANUAL REMEDIATION

Humans stuck bridging the gap

Developers manually copy error messages, craft new prompts explaining the failure context, and hope the agent fixes it without breaking something else.

Workflow

The closed loop in 4 steps

The AI that wrote your code hands off intent to the AI that tests it.

1

Code naturally

Use Claude Code as you always do. The bridge hook captures each edit — no extra workflow steps.

2

Intent Capture

The local bridge daemon captures files touched, prompt summaries, and agent reasoning, and streams them to the platform.

3

Playwright Sandbox

Spins up a preview container, logs in with test credentials, and exercises user journeys.

4

One-Click Fix

Failures yield exact remediation markdown prompts. Paste back into your agent to resolve.

Forensic Artifacts

Forensic proof for every run

Recordings and traces captured together for every run, so nobody has to reproduce failures by hand.

remediation-prompt.md
## 🚨 Autonomous PR Verification Failed: checkout
> **Risk Assessment:** HIGH RISK
> **Affected Surfaces:** /checkout, /dashboard

### 🔍 Forensic Evidence
- **Error:** 'Invalid customer address for Apple Pay token invoice generation'
- **Video:** artifacts/runs/feat-quick-checkout_f1e2d3c4/video.webm
- **Trace:** artifacts/runs/feat-quick-checkout_f1e2d3c4/trace.zip

### 🧠 Intent Context Hand-off
- User Prompt: "Add instant Apple Pay button to checkout"
- Agent Intent: "Bypass standard address form for 1-tap checkout"
- Touched: components/pay-button.tsx, app/checkout/page.tsx

### 📋 Ready-to-Paste Remediation Prompt for Claude Code:
Fix regression in /checkout: Ensure the Apple Pay session handler passes the default billing address token downstream to the invoice service.
Comparison

Traditional CI vs. Autonomous QA Engine

CapabilityTraditional CI SuiteRazeQA Platform
Test GenerationOnly what engineers hardcodedIntent-driven user journeys
Adjacent Bug DetectionRegressions slip through to prodAI analyzer flags downstream routes
Forensic ProofPlain text terminal logsSession video + Playwright trace
RemediationManual reproducing & debugging1-click prompt for Claude / Cursor
Freshness CachingRe-runs identical tests on pushSHA dedup returns the cached run on re-check
System Architecture

Isolated sandboxes with a clear data-retention story

Built for teams that need per-run isolation, auditable artifacts, and a retention window they control.

Autonomous QA Execution Pipeline
Disposable container per run
STAGE 01

Local Bridge Daemon

Hooks into the Claude Code tool loop. Streams captured intent, prompt summaries, and the files touched to the platform.

STAGE 02

Intent & SHA Cache

Indexes branch sha, touches Supabase intent logs, and validates freshness against existing baselines.

STAGE 03

Ephemeral Sandbox

Boots a disposable container for the app under test, injects test credentials, and drives it with a headless browser that records video and trace.

STAGE 04

PR Forensics & Fix

Posts native GitHub Check Run, embeds forensic video link, and generates instant fix prompt for your AI.

Disposable Runtime

The pull request is cloned inside a "--rm" container and provisioned there; the container is destroyed when the run ends. Forensic artifacts (video, trace, screenshots, redacted DOM) are stored in a private bucket behind short-lived signed URLs and removed on a configurable retention window.

AES-256-GCM Credential Vault

Test-user credentials are encrypted at rest with AES-256-GCM. The key is supplied out-of-band via CREDENTIAL_STORE_KEY; credentials are injected into the sandbox as environment variables at boot and are never written to the repo, the intent log, or an LLM prompt.

Resource-Limited Sandboxes

Each run gets a container capped on memory, CPU and PIDs, with dropped capabilities and no privilege escalation. The headless browser runs from the engine against that container, and session state is reused across journeys within a single run for deterministic replay.

Catch regressions before your users do.

Connect the bridge daemon to your repo in 60 seconds. Start getting forensic proof and one-click fixes for every commit.