Settings

Theme

Agent runtime reduces LLM turns by 80% with a higher success rate in DeepSWE

github.com

2 points by yohji1984 10 days ago · 1 comment

Reader

yohji1984OP 10 days ago

Hi HN, I've been working on an agentic runtime framework. The earlier benchmark eval is promising. But I can see the limits of the test design and the fragility of the runtime itself. I would like to ask for your reviews of the framework and the eval process itself.

https://github.com/Tura-AI/tura

https://turaai.net/docs#benchmark-current-test-set-record

Tura-AI/tura https://turaai.net/blog#why-i-am-building-tura

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection