Settings

Theme

Show HN: Darts Vision Benchmark

darteval.vercel.app

1 points by red545 · 0 comments · 1 min read

Reader

Detecting the correct score from a dartboard photo is unexpectedly hard for LLMs.

The task appears to stress spatial reasoning: Gemini 3 models lead this benchmark by a decent margin.

Counterintuitively, “more reasoning” often reduces accuracy.

Even the top-performing model scores only ~36% of darts correctly.

No comments yet.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection