Agent runtime reduces LLM turns by 80% with a higher success rate in DeepSWE
github.comHi HN, I've been working on an agentic runtime framework. The earlier benchmark eval is promising. But I can see the limits of the test design and the fragility of the runtime itself. I would like to ask for your reviews of the framework and the eval process itself.
https://github.com/Tura-AI/tura
https://turaai.net/docs#benchmark-current-test-set-record
Tura-AI/tura https://turaai.net/blog#why-i-am-building-tura