Live data from Hacker News

Show HN: LLMs suck at writing integration code… for now

github.com

11–15 of 15 posts

Re: Show HN: LLMs suck at writing integration code… for now

#11
Awesome work Stefan, this is super insightful! Really appreciate the transparency and open-sourcing the benchmark. The 68% success rate is a wake-up call for anyone building with LLMs. Your 91% integration layer result is impressive, shows tooling matters. Excited to see what you uncover next with MCP!

Re: Show HN: LLMs suck at writing integration code… for now

#13
Really appreciate you sharing this. What I am trying to use is gpt o3, so would be curious to see it in the benchmarks. Still seeing the raw traces tells me the tooling is starting to cross the “actually usable” line and makes me want to try on my examples this weekend. Looking forward to the MCP benchmark as well.
Post reply on HN