We just open-sourced Flightplanner — a framework for spec-driven E2E testing in the age of AI agents.
Here's the problem: AI agents made writing code cheap. But maintaining it? That's a different story. E2E tests are the worst offenders — expensive to write, slow to run, painful to maintain. Yet they're the only tests that validate your product from the user's perspective.
Nobody celebrates a well-maintained E2E suite. But everyone notices when it's missing.
Our insight: what if instead of maintaining brittle test code, you maintained a human-readable spec of how your product should behave?
A Flightplanner spec is three things at once:
- Documentation — tells anyone on the team what the product does
- A product contract — bridges product thinking and engineering, like user stories but inside the codebase
- A testable artifact — an agent reads it and generates the actual test code
You write plain-language specs like:
- "Displays the dashboard after login"
- "Shows the user's name in the navigation bar"
Each bullet is a concrete, verifiable assertion. The spec is the single source of truth — when specs and tests disagree, the spec wins.
The intent stays human. The implementation becomes automated.
When a test breaks, you trace it back to a specific behavior in plain language — not reverse-engineering what "expect(page.locator('.dashboard-header')).toBeVisible()" was supposed to mean.
Flightplanner provides a growing set of skills for agents. Some of the core ones:
- fp-init — bootstrap specs by analyzing your source code
- fp-update — generate or sync tests when specs change
- fp-fix — fix failing tests (never touches the specs)
- fp-audit — find gaps between your codebase and specs
It works with any test framework (Playwright, Cypress, pytest…) and any AI agent.
Get started:
We've been using it on our own projects and it's changed how we think about E2E testing. We hope it does the same for you.
github.com/endorhq/flightplanner
Happy testing!