AI & AutomationFeb 20268 min read

Building an AI-Powered Test Case Generator: What GhostAI Taught Me

GhostAI crawls a web app, parses the DOM, and auto-generates test cases — no scripting required. Building it taught me where AI genuinely helps QA, where it embarrasses itself, and why the crawler is the hard part.

DILJOT SINGH · TECHNICAL LEAD

As Lead Infrastructure & Backend Architect on an AI-powered QA platform, I built GhostAI: a crawler that explores a web application, parses its DOM, and generates executable test cases automatically. The pitch is seductive — point it at your app, get a regression suite. The reality is more interesting, and the lessons generalize to most "AI will automate X" projects.

The crawler is 70% of the problem

Before any AI touches anything, you need a faithful map of the application — every reachable state, every interactive element, every form and its validation behavior. Modern web apps make this genuinely hard: content renders asynchronously, routes are client-side, elements appear on hover, and half the interesting states live behind multi-step flows with authentication.

Diagram
Fig 1 — GhostAI crawl-to-test pipeline: crawler → state model → generator → containerized runners

Generating tests people can read

The naive output — a click-path replay with hard-coded selectors — is worthless. It breaks on the first UI change, and no human can tell what it was verifying. Two design decisions fixed this. First, generated tests target semantic anchors (roles, labels, stable data attributes) with a fallback chain, so a CSS refactor doesn't zero the suite. Second, every generated test states an intent — "submitting the signup form with an invalid email shows a validation error" — derived from the DOM patterns around the interaction (form semantics, validation attributes, error-message containers). Intent is what makes a generated test reviewable, and reviewability is what makes it trustworthy.

Where the AI actually earns its keep

Honest scorecard from production use: the AI is genuinely good at coverage breadth (it explores paths human testers skip out of boredom), at generating plausible form inputs (including the boundary cases humans forget), and at flagging "this state looks anomalous" for human triage. It is mediocre-to-embarrassing at knowing what matters — it will test a footer link with the same enthusiasm as the checkout flow — and at asserting business correctness it has no way to know (is a $0.00 total a bug or a free tier?).

The honest framing isn't "AI writes your tests." It's "AI drafts coverage; humans supply judgment." Products that promise the first and deliver the second erode trust. We repositioned the UX around review-and-promote, and adoption went up.

Demo — GhostAI exploring an app and drafting test cases (3 min)

Infrastructure: execution is a scheduling problem

The platform also ran user-authored Cypress, Playwright, and JMeter suites, which meant the execution layer had to treat test runs as untrusted, resource-hungry workloads: each run in its own container with CPU/memory limits and no network path to other tenants, queued through the same SLA-separated async patterns I use everywhere else, with artifacts (videos, traces, HAR files) shipped to object storage. Generated tests compiled down to the same execution format as human-written ones — one runner, one artifact pipeline, no special cases.

Takeaways