Live data from Hacker News

GPT-5

openai.com

261–270 of 1001 posts

Re: GPT-5

#261
The incremental improvement reminds me of iPhone releases still impressive, but feels like we’re in the ‘refinement era’ of LLMs until another real breakthrough.

Re: GPT-5

#262
post #51

The silent victory here is this seems like it is being built to be faster and cheaper than o3 while presenting a reasonable jump, which is an important jump in scaling law On the other hand if it's just getting bigger and slower it's not a good sign for LLMs

Yeah, this very much feels like "we have made a more efficient/scalable model and we're selling it as the new shiny but it's really just an internal optimization to reduce cost"

Re: GPT-5

#264
post #113

> With GPT-5 we will be deprecating all of our prior models Wow, they actually did it

GPT-5 is likely much cheaper to serve, and that's the "big win" here, not necessarily any improvement in output.

Re: GPT-5

#266

74.9 SWEBench. This increases the SOTA by a whole .4%. Although the pricing is great, it doesn't seem like OpenAI found a giant breakthrough yet like o1 or Claude 3.5 Sonnet

I'm pretty sure 3.5 sonnet always benchmarked poorly, despite it being the clear programming winner of it's time.

Re: GPT-5

#267

In terms of raw prose quality, I'm not convinced GPT-5 sounds "less like AI" or "more like a friend". Just count the number of em-dashes. It's become something of a LLM shibboleth.

I've worked on this problem for a year and I don't think you get meaningfully better at this without making it as much of a focus as frontier labs make coding.

They're all working on subjective improvements, but for example, none of them would develop and deploy a sampler that makes models 50% worse at coding but 50% less likely to use purple prose.

(And unlike the early days where better coding meant better everything, more of the gains are coming from very specific post-training that transfers less, and even harms performance there)

Re: GPT-5

#268
Is this good for competitors because it's so underwhelming, or bad for AI because the exponential curve is turning sigmoid?

Re: GPT-5

#270

Someone at OpenAI screwed up the SWE-bench graph. o3 and GPT-4o bars are same height, but with different values.

The graph is more screwed up than that: the split bar is also split in a nonsensical way

It feels a bit intentional

Post reply on HN