Awesome work Stefan, this is super insightful!
Really appreciate the transparency and open-sourcing the benchmark.
The 68% success rate is a wake-up call for anyone building with LLMs.
Your 91% integration layer result is impressive, shows tooling matters.
Excited to see what you uncover next with MCP!
Show HN: LLMs suck at writing integration code… for now
11–15 of 15 posts
Re: Show HN: LLMs suck at writing integration code… for now
#12Exciting benchmarks, great work Adina and Stefan!
Re: Show HN: LLMs suck at writing integration code… for now
#13Really appreciate you sharing this. What I am trying to use is gpt o3, so would be curious to see it in the benchmarks. Still seeing the raw traces tells me the tooling is starting to cross the “actually usable” line and makes me want to try on my examples this weekend. Looking forward to the MCP benchmark as well.
Re: Show HN: LLMs suck at writing integration code… for now
#14Thanks for the self host option.
I tried the slack example and was very impressed with results, thank you!
Re: Show HN: LLMs suck at writing integration code… for now
#15[dead]