I Had Claude and Codex Rewrite the Same App. The One With Better Architecture Lost.

· Medium ·

1 min read Original article ↗

Illya Yalovoy

One human, one month, two AI fleets, 270 commits, and 2,243 passing tests — measuring the resource that coding benchmarks ignore: human attention.

Press enter or click to view image in full size

Last month I ran an instrumented natural experiment: I had two frontier AI coding systems independently reimplement the same real application from the same reference codebase, using the same task-orchestration tool and supervised by the same increasingly impatient human — me.

The application is Axiotask, a Google Tasks client I originally built, using Rust, Tauri, and Svelte.

Both rewrites used Dart and Flutter:

  • Rewrite #1: Claude Code, primarily Opus 4.8 workers, with Fable 5 involved in planning and Opus 5 used for the final tasks.
  • Rewrite #2: OpenAI Codex, using GPT-5.6 sol and later terra.

Both produced working, analyzer-clean applications with green test suites. Both were genuinely impressive.

Codex produced the more rigorously specified and aggressively tested synchronization engine. Its desired-state architecture, process-death recovery, and evidence-linked test matrix are stronger than the corresponding designs in my original Rust application.