AI agents produced 212 merged PRs and 111,000+ changed lines in 30 days — but human review and machine verification failed to keep pace.
Speed was never the constraint — verification was. 25–30% of all PRs were rework; stabilization now precedes every new feature.
Review is a budget, not a virtue. One human can properly review ~30 high-risk PRs a month, not 212. Risk-tiered sign-off replaces reviewing everything (badly).
Self-reported green is not green. Agents’ own success claims went uncorroborated for weeks — broken CI and skipped deploys stayed “green” for 16 days. Verification must be machine-enforced and independent.
Honesty needs a consumer. Agents disclosed their shortcuts in every PR; nobody acted. Disclosed debt now auto-converts into tracked work.
On July 11th, at 07:54:14 UTC, an AI agent opened a pull request against my repository: 9,397 lines added, 478 removed, 67 files changed. The PR body was two words: “version update.”
It merged at 07:54:22. Eight seconds later.
No CI ran on it. No human read it. I know, because the human was me.
This is a midway evaluation of an experiment I’ve been running for the past month: building a real, multi-tenant SaaS product almost entirely with AI coding agents, under a workflow I designed to keep them honest. The product isn’t finished. But the data from the first 30 days — 212 merged PRs, over 220 completed issues, 111,599 lines changed — taught me something more valuable than any feature I shipped:
When AI agents can produce work faster than you can verify it, every human-paced control in your process silently degrades to zero. And the system keeps reporting green the whole way down.
The product is a multi-tenant agentic SaaS platform for small-business owners: a place to hire digital workers. A business owner browses a library of AI employees — a quoter, a lead generator, an invoicing clerk, a dunning agent — hires one, connects their real business systems (Zoho Books, Google Workspace) through external connectors, and the digital worker does the job: drafts quotations from a real price list, chases unpaid invoices, generates UAE-compliant e-invoices, finds leads.
The stack is deliberately boring and production-shaped: a Bun + Turborepo monorepo, Hono API, Next.js 15, Drizzle/Postgres with pgvector, and Terraform on GCP — Cloud Run, Cloud SQL, Secret Manager, keyless deploys via Workload Identity Federation.
The delivery workflow is the interesting part. I defined a loop tightly coupled to GitHub: I write plan files; a sync process turns them into issues; AI agents claim issues via server-side branch refs, build to acceptance criteria, run a ladder of deterministic local gates, and open one PR per issue. Humans do exactly two things in this design: write plans and review PRs.
The agents held up their end. I didn’t hold up mine. That’s the story.
212 merged PRs (July 3–25), 111,599 lines changed
571 commits — all but a couple authored by agents
~19.6k lines of source code, ~11.4k lines of test code
Whole platform skeleton — monorepo, DB, API, web app, agents service, billing, notifications, CI — merged in the first 22 hours
Nine UI screen PRs (~4,700 lines) merged in 39 minutes
Peak day: 59 merges on July 4th
And the numbers I’m less proud of:
208 of 212 PRs have zero code review. Median time-to-merge: about 90 seconds. Nineteen PRs merged in under a minute.
~25–30% of all PRs were rework — fixes, remediations, or re-plans of work merged days or hours earlier.
The CI gates job was broken for 16 days — and PRs kept merging over the red check the entire time.
The staging environment, which every dashboard showed as deployed, was serving Terraform’s placeholder “hello” image the whole month.
Before the autopsy, fairness. The agent-written code is genuinely better than its reputation deserves.
Every implementation PR mapped acceptance criteria one-to-one to named tests. Several PRs mutation-checked their own tests — deliberately breaking the fix to confirm the test goes red. The database tests run against in-memory Postgres with real semantics, not mocks. The gate ladder grew from 6 to 12 deterministic checks, and some of them are ideas I’d carry to any team: a boundaries gate enforcing import direction, an identity gate that statically forbids an entire class of tenancy bug, a config gate that blocks committed cloud project IDs, a run-secrets gate asserting every runtime env var is mounted from Secret Manager.
The security engineering, on paper, is real: an AES-256-GCM credential vault with per-value IVs, HMAC-signed OAuth state, constant-time comparisons, deterministic integer-based money math with tests asserting a hostile LLM cannot alter stored amounts.
And the agents were honest. Relentlessly, uncomfortably honest. They flagged their own shortcuts in code comments. They wrote “Honest gaps” sections in PR bodies. One agent shipped a security-relevant infrastructure change with a section literally titled:
Two deliberate calls, worth a look during review
Cloud Run invocation is public (
allUsersinvoker), ingress unrestricted. … Say the word if you’d rather pay for a connector/Direct VPC and lock ingress down instead.
It asked for review. I merged it 56 seconds later.
The agents did what they were incentivised to do — ship plausible, self-certified work fast. So did I. Neither of us was the verification layer.
The verification vacuum
During the analysis I found that the reworks and near-misses all trace to one systemic condition: a verification vacuum. There was a human verification step in the process — it just wasn’t followed diligently, and everything downstream inherited that. Three failures, compounding:
Every PR body carried a gate table — format ✅, typecheck ✅, test ✅ — self-reported by the same agent that wrote the code. Plausible, usually true, and corroborated by nothing. The system optimized for the throughput of narratives about working software rather than the deployment of working software.
Here’s how far that gap stretched. On July 9th, a framework PR (#179) removed the Bun installation step from the CI gates workflow. From that moment, every CI gates run died at bun: not found — exit 127 — before executing a single check. The fix landed July 25th, sixteen days later. In between, roughly 25 PRs merged with a red gates check.
The PR that finally fixed it described the damage better than I can:
Gate 1 therefore died at
bun: not found(exit 127) before a single gate executed, so the reviewer’s machine-verified backstop verified nothing and trained reviewers to read a redgatescheck as noise.
And deployments? Same story, worse punchline. The deploy jobs evaluated their enable-flag before GitHub’s environment context existed, so the flag always read empty and the jobs silently skipped — on every push to main, for weeks, while runs reported green:
…both jobs were skipped on every push to main while the run reported green. The staging Cloud Run services still serve Terraform’s placeholder
helloimage on revision 00001.
Green dashboards. Twelve gates. Five hundred and seventy commits. And a staging environment running “hello”.
I have to own this one. 208 of 212 PRs received no review. I scrubbed through some of the code; mostly I merged — some PRs in under 60 seconds, one in 8 seconds. The workflow’s design said humans do two things: write plans and review PRs. Under the pressure of agent-speed throughput, review quietly collapsed into ratification, and nobody — including me — ever made that decision explicitly. It just happened.
The cost wasn’t abstract. The self-flagged public gateway merged unread. A schema storing per-tenant API keys in plaintext merged in four minutes, carrying its own confession in a comment:
// ponytail: virtual_key is a secret in plaintext; move to a secret manager or
// column encryption before real tenant traffic.
export const tenantGatewayTeams = pgTable("tenant_gateway_teams", {
tenantId: text("tenant_id").primaryKey(),
teamId: text("team_id").notNull(),
virtualKey: text("virtual_key").notNull(),That comment is still in the schema three weeks later. A hand-rolled Stripe webhook verifier — competent enough to use constant-time comparison, but missing the replay-window check the official SDK enforces — has survived since day two. Not because anyone judged these acceptable. Because no one judged them at all.
The workflow includes a hard scope rule: ~400 changed lines or ~6 files per PR — over that, stop and split. 62 of 212 PRs broke it. Not by sneaking past it — by negotiating with it, in prose, in the PR body:
Precedent for merging over the cap exists (the drizzle migration PRs, #251). If you’d like the gate to stop flagging infra work, the fix is a
size_exclude_pathsentry…
One PR merged over a knowingly red size check “at the owner’s explicit direction” — my direction. Another admitted to choosing worse code placement specifically to duck the file-count limit. I also manually bypassed some of the non-trivial checks myself when they were inconvenient. Every bypass created precedent for the next one.
That’s the lesson in one sentence: a control you can argue with in a PR body is not a control. The three worst-outcome PRs of the bootstrap phase — a trunk corruption, unreviewed payment code, a feature that spawned eight follow-up fixes — were exactly the three biggest PRs.
There’s one more bill that came due, and it deserves its own section because it’s the most visible consequence of building UI-first on mocks with no reviewer asking questions.
On July 24th, the agents ran a systematic remediation sweep of the product’s own interface — roughly ten PRs whose only job was deleting fiction that earlier agent phases had shipped. The PR bodies read like a confession log:
The library worker page rendered every catalog integration with a hardcoded green check, so
/library/quoterclaimed it works with Gmail (no connector at all)… The section was decoration dressed as status.Removes the hire flow’s “Connect a data source” step — a simulated destinations list where two of three providers always failed server-side (theater, not function).
Home feed stats were computed-looking constants (
leads: 0,approvalRate: 1,hoursSaved: 0)…
A constant 100% approval rate. Invented customer activity presented as the tenant’s own. Save buttons that saved to nothing. A hardcoded $199/month sitting next to a real Stripe checkout. Each of these would have died in thirty seconds of genuine review with one question: does this button do anything?
The conclusion I’ve landed on for the second half of this build: truth is not a negotiated statement. It must be hardened at the infrastructure level, where neither the agent nor — frankly — I can bypass it in a moment of throughput enthusiasm. The current risk profile is unacceptable, and stabilization comes before any new feature.
The burden of proof moves from the author to the infrastructure. Three layers.
Branch protection: red = unmergeable. No exceptions, no override culture. CI gates run on every PR to main — including framework branches and side branches, which previously skipped CI entirely (the 9,397-line PR had no CI at all, by branch filter).
Automated failure alerting on main. Both multi-week silent outages — dead CI, skipped deploys — were one alert away from being one-day incidents.
The P0 security list, fixed through the pipeline: real Firestore rules (the current ones are allow-all with an expiry date), the Stripe replay window, encrypting the plaintext tenant keys with the vault that already exists in the same codebase.
Strict enforcement directive: no out-of-band fixes. Every remediation lands through CI/CD, versioned and reproducible. Manual console surgery is how the environment drifted from the code in the first place.
The strategic frame: eliminate the gap between code claims and deployed reality by institutionalizing machine-led skepticism.
Tiered review replaces zero review:
Automated adversarial review — every PR gets a pass from a dedicated reviewer agent whose job is to refute the work, with a recorded, binary verdict. Not a second opinion; a prosecution.
Risk-based human sign-off — mandatory human review for every High-Risk PR: schema, auth, secrets, billing, infrastructure. That was ~30 PRs this month. One person can review 30 PRs a month properly. One person cannot review 212 — and pretending otherwise is how you get a merge button.
Minimum merge delays — 15 minutes on high-risk changes. Its only purpose is making 8-second merges structurally impossible.
Integration gates — CI boots the API and agents against dockerized Postgres and runs the migrations, ending the environment-parity sagas where tests passed on in-memory databases while every real environment drifted. Migration collisions (parallel agents claimed the same migration number three separate times this month) get resolved structurally: timestamp-based IDs or a journal-sequence check that fails the second claimant.
Stability is the primary driver of speed. The 25–30% rework tax is the capacity I’ve been missing; most of it traces upstream to plans that never said how success would be proven, what resources they locked, or what they were deferring. Four invariants, enforced at plan-sync time, with a Risk Tier assigned in every plan header:
Resource locking. Plans declare exclusive resources — migration sequence, specific routes,
schema.ts. Two plans claiming the same resource serialize; they don’t collide in a merge queue three PRs later.Verifiability rule. Every acceptance criterion ships with a proof command. “UI looks correct” is banned; a Playwright behavior test or it doesn’t count. (This month, UI criteria were “verified” by tests asserting the source code contained certain strings — tests that pass while the interface is broken.)
Vertical-slice rule. No plan ships a mock or stub as its deliverable. If a stub is unavoidable, the repayment plan is created in the same planning PR — linked, scheduled, non-optional. The de-fake tax gets paid at planning time, once, instead of at remediation time, ten PRs at a time.
Size budgeting at plan time. The 400-line budget (excluding generated files — which inflated every migration PR and gave every author a standing excuse) is estimated and enforced before code exists. The split mechanism demonstrably works when it’s actually forced; it just fired after the fact.
Debt ingestion closes the honesty loop. An automated scanner harvests every TODO-marker convention: marker and every “deferred scope” declaration from merged PRs and converts them into planning items for the next batch. The agents’ honesty finally gets a consumer. Disclosure without follow-through stops being an option — for them and for me.
Self-reported green is not green. Every silent failure this month lived in the gap between an agent’s claim and independent corroboration. Close that gap with machines, because you will not close it with attention.
Your controls must be machine-enforced or they are decorative. Every deterministic gate held. Every social control — size limits, “human reviews before merge” — was negotiated away within days, by the agents and by me.
Deploy from day two, and verify the deployment, not the dashboard. The most expensive lie in my repo wasn’t in any diff. It was a green checkmark on a deploy job that never ran.
Agent honesty is abundant and worthless without a consumer. Build the pipe that turns disclosures into work items, or watch them fossilize into the exact debt they warned you about.
Review capacity is a budget, not a virtue. Decide what gets human eyes by risk tier, in advance — or throughput will decide for you, and it will decide zero.
The experiment continues. The agents write the code; the infrastructure now writes the truth. I’ll report back when the smoke test goes green on a staging environment that is, finally, running something other than “hello”.
Numbers computed from the GitHub API across all 212 merged PRs (July 3–25). Code excerpts and PR quotations are verbatim from the repository at the time of writing.








