Live data from Hacker News

Agent runtime reduces LLM turns by 80% with a higher success rate in DeepSWE

github.com

1–2 of 2 posts

Re: Agent runtime reduces LLM turns by 80% with a higher success rate in DeepSWE

#2
Hi HN, I've been working on an agentic runtime framework. The earlier benchmark eval is promising. But I can see the limits of the test design and the fragility of the runtime itself. I would like to ask for your reviews of the framework and the eval process itself.

https://github.com/Tura-AI/tura

https://turaai.net/docs#benchmark-current-test-set-record

Tura-AI/tura https://turaai.net/blog#why-i-am-building-tura