Settings

Theme

Atlas-Finance: Evaluating AI Agents Inside a Bank

joinhandshake.com

2 points by cjbarber · 1 comment

Reader

1 thread
cjbarberOP

I found this interesting.

> We evaluate 11 frontier models (see Figure 1), all run within the OpenCode agentic harness. Claude Opus 5 performs the best, yet still only manages to pass 12.3% of tasks. Claude Fable 5.1 and GPT-6 Astra are close behind, but the other eight models have significantly lower pass rates.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection