Live data from Hacker News

GPT5 is worse than 4.1-mini for text and worse than Sonnet 4 for coding

news.ycombinator.com

1–10 of 17 posts

GPT5 is worse than 4.1-mini for text and worse than Sonnet 4 for coding

#1
It seems that OpenAI have got the PR machine working amazingly. The Cursor CEO said it's the best, as did Simon Willison (https://simonwillison.net/2025/Aug/7/gpt-5/).

But I've found it terrible. For coding (in Cursor), it's slow, fails with tool calls often (no MCP just stock Cursor tools) and stored some new application state in globalThis - something that no model has ever attempted to do in over a year of very heavy Cursor / Claude Code use).

For a summarization/insights API that I work on, it was way worse than gpt-4.1-mini. I tried both mini and full gpt5, with different reasoning settings. It didn't follow instructions, and output was worse across all my evals, even after heavy prompt adjustment. I did a lot of sampling and the results were objectively bad.

Am I the only one? Has anyone seen actual real-world benefits of GPT-5 vs other models?

Re: GPT5 is worse than 4.1-mini for text and worse than Sonnet 4 for coding

#5

it solved a huge bug i've been struggling with.

Had Sonnet 4 not been able to?

No, it kept going in circles....spent like 3 weeks trying to fix it. Got access to gpt5 yesterday and all major bugs are resolved.

Re: GPT5 is worse than 4.1-mini for text and worse than Sonnet 4 for coding

#8
GPT-5 isn’t really a brand-new model in the way people think. From what I’ve seen, the goal was more about reducing costs and unifying the interface than releasing a totally different architecture. Under the hood it is still routing to models we already know, just picking what it thinks will give the “best” result for the request.

That can be fine for a lot of general use cases, but if you’re working in specific domains like coding agents or high-precision summarization, that routing can actually make results worse compared to sticking with a model you know performs well for your workload.

Re: GPT5 is worse than 4.1-mini for text and worse than Sonnet 4 for coding

#9
I feel like they should have let GPT 5 overlap in experimental mode for a month or so. It took a while to get the kinks out of GPT-4 until people trusted it. Just switching it on is really hurting their brand.

The fact they didn’t do this makes me think their finances are in very bad shape.

Re: GPT5 is worse than 4.1-mini for text and worse than Sonnet 4 for coding

#10

Earlier quoted context omitted.

Had Sonnet 4 not been able to?

No, it kept going in circles....spent like 3 weeks trying to fix it. Got access to gpt5 yesterday and all major bugs are resolved.

Interesting I tried it to fix some unit tests that were failing but made the problem worse. Sonnet was able to fix the failing unit tests and the new problems introduced by GPT5. I used Claude Code for Sonnet and Cursor Agent for GPT-5. Maybe Cursor Agent is just bad?
Post reply on HN