Settings

Theme

bisonbear

Karma
63
Created
1 year ago

About

Building evals for AI coding agents, on your repo. Tests pass. Nobody's measuring the rest. http://stet.sh email ben@stet.sh

Recent Submissions

  1. 1. I compared Opus 4.8 vs. Opus 5 on 25 of my tasks to see what the difference was (stet.sh)
  2. 2. I compared 5 popular token saving methods in Codex and found that none delivered (stet.sh)
  3. 3. I ran Sonnet 5 vs. Opus 4.8 head to head on 24 tasks to see what's different (stet.sh)
  4. 4. I evaluated GLM 5.2 against the frontier on tasks from real repos (stet.sh)
  5. 5. I benchmarked Opus 4.8 vs. GPT 5.5 on 2 open source repos (stet.sh)
  6. 6. I used autoresearch to improve my AGENTS.md, measured against real tasks (stet.sh)
  7. 7. A brief investigation into the GPT-5.5 regression claims (stet.sh)

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection