Autonomous QA Agent

What if your full regression
suite built itself
overnight?

Achilles is an autonomous quality assurance solution that helps QA engineers build test automation and run hands-on testing at an accelerated rate. One sentence kicks off a complete eight-phase pipeline — no babysitting required.

$ npm i @civitas-cerebrum/achilles

Everything your QA team needs,
fully autonomous

Achilles routes each request to the right working mode. Give it a task in plain English; the harness takes care of the rest.

🚀

Zero-to-Suite Onboarding

One command turns a project with zero tests into a fully maintained suite. An eight-phase autonomous pipeline handles scaffolding, journey mapping, coverage expansion, adversarial bug discovery, and delivers a branded stakeholder summary deck.

Onboarding
📈

Coverage Expansion

Grows the suite journey by journey, running independent journeys in parallel. Three compositional passes and two adversarial passes deliver rigorous coverage — and every state-changing step is proven by an API or database oracle, not just a toast. Depth mode enforces strict per-journey parallelism for high-stakes audits.

Expansion
🐛

Adversarial Bug Discovery

Probes the live app fresh — before reading any context — to catch what familiarity blinds you to. Returns a deduplicated bug ledger where every finding is evidence-backed, ranked by severity and business priority, and tracked through a triage lifecycle — with reproduction tests, IDOR probes, race conditions, and state-skip vulnerabilities.

Bug Discovery
🔧

Automated Suite Repair

Batch-clusters failures by shared root cause and heals entire clusters at once. The endless maintenance cycle, automated. Stops the slow bleed of a degrading test suite before it reaches your CI pipeline.

Suite Repair
🔍

Companion Mode

On-demand, evidence-first verification of a single flow. Returns a shareable bundle: per-step screenshots, video recording, Playwright trace, HAR file, console log, and a pass/fail summary — ready to drop into any ticket.

Companion
🤖

AI Safety Testing

Red-team the LLM features in your app. A three-agent architecture — adversary, target, and judge — tests guardrails, prompt injection resistance, bias detection, and compliance. Produces reproducible adversarial findings.

Agents vs Agents

Eight phases. One command.
Fully autonomous.

The onboarding pipeline runs end-to-end without intermediate confirmations. Harness hooks gate every phase transition — the agent finishes the contract or surfaces a blocker for human triage.

01 —
Scaffold
Playwright framework, test directories, and config set up from scratch.
02 —
Groundwork
App is crawled to discover all pages, routes, and interactive surfaces.
03 —
Happy Paths
Primary user journeys are automated first — the critical happy-path layer.
04 —
Journey Map
All discoverable user journeys mapped and prioritised by business impact (P0–P3).
05 —
Coverage
Multi-tiered compositional and adversarial passes expand coverage across every journey.
06 —
Bug Hunt
Adversarial probing of the live app returns a deduplicated, prioritised bug ledger.
07 —
Secrets Sweep
Hardcoded credentials, API keys, PII, and URLs extracted into env variables.
08 —
Summary Deck
Branded HTML/PDF work summary delivered for stakeholders. Done.

You stay in the loop where it matters: reviewing, prioritising, deciding. Achilles handles everything else.

Civitas Cerebrum · Achilles Design Principle

Built different,
by design

Achilles isn't a test generator — it's a full autonomous QA methodology with deterministic phase enforcement and structured data contracts baked in.

Harness-Enforced Discipline

20+ hook scripts (23+ registered gates) sit between every tool invocation. They prevent scope compression, phase skipping, and silent scope reduction — the failure modes that plague AI-assisted testing.

🎯

Senior-Grade Judgment

Probes run fresh — fresh eyes catch what familiarity blinds you to — then findings are risk-weighted by defect likelihood, ranked by severity and business priority, and held to an oracle stronger than a passing toast. The calls a senior QA engineer makes, not just green or red.

📐

Declarative Steps API

Tests reference elements by name, never raw selectors. Selectors live in a page repository and are validated against the live DOM before any test runs — zero flakiness from stale locators.

📦

Structured Data Contracts

Every schema-bearing dispatch is gated against a JSON-Schema return contract — cited in the brief before dispatch, validated on return. Composer, probe, and reviewer returns carry a typed handover envelope for deterministic phase-to-phase handovers.

🔄

Autonomous CI/CD Mode

External CLI drivers can invoke Achilles hands-off for scheduled nightly QA pipelines. The pipeline finishes the contract or surfaces a blocker — it never silently truncates.

📊

Stakeholder-Ready Output

Not just test code. Journey maps, test catalogues, bug ledgers, and branded summary decks communicate QA value to product owners and executives — automatically.

Built on

Playwright JSON Schema Pixelmatch Babel Parser

Ready to automate
everything?

One install. One sentence. A complete test suite by morning.

MIT License Built on Playwright Driven by Agents v0.1.0