noemon (@noemon_ai) on X

X (formerly Twitter) ·

1 min read Original article ↗

Announcing: ARC-AGI-2 top score and at a fraction of the cost through an agentic learning harness for iterative self-improvement. Public Eval — 91% @ $3.3/task. SOTA compared to even new models trained for extreme reasoning like GPT-5.4 Pro and Gemini 3 Deep Think, while also reducing their cost by more than 4x. We used Gemini 3.1 Pro (Preview) and we increased its score by 12 percentage points.

@GregKamradt@fchollet@mikeknoop