Settings

Theme

Show HN: Gemini models are getting better – we can prove it

github.com

2 points by favurdev · 0 comments · 1 min read

Reader

We ran seven Google models through the same task, in the same multi-agent harness.

Improvements to reasoning and tool calls are the main reasons for the improvements. We saw less and less reasoning and some of the lowest invalid tool calls across the entire evaluation suite.

Github link for the run output and analysis. More at evals.favur.dev

No comments yet.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection