One human, one month, two AI fleets, 270 commits, and 2,243 passing tests — measuring the resource that coding benchmarks ignore: human attention.
Press enter or click to view image in full size
Last month I ran an instrumented natural experiment: I had two frontier AI coding systems independently reimplement the same real application from the same reference codebase, using the same task-orchestration tool and supervised by the same increasingly impatient human — me.
The application is Axiotask, a Google Tasks client I originally built, using Rust, Tauri, and Svelte.
Both rewrites used Dart and Flutter:
- Rewrite #1: Claude Code, primarily Opus 4.8 workers, with Fable 5 involved in planning and Opus 5 used for the final tasks.
- Rewrite #2: OpenAI Codex, using GPT-5.6 sol and later terra.
Both produced working, analyzer-clean applications with green test suites. Both were genuinely impressive.
Codex produced the more rigorously specified and aggressively tested synchronization engine. Its desired-state architecture, process-death recovery, and evidence-linked test matrix are stronger than the corresponding designs in my original Rust application.