Live data from Hacker News

GPT-5

openai.com

911–920 of 1001 posts

Re: GPT-5

#911
post #896

Anecdotal review: Been using it all morning. Had to switch back to 4. 5 has all of the problems that 2/3 had with ignoring any context, flagrantly ignoring the 'spirit' of my requests, and talking to me like I'm a little baby. Not to mention almost all of my prompts result in a several minute wait with "thinking longer about the answer".

Yea I see this a lot with Gemini since 2.5

Very stubborn and “opinionated”

I think most models will tend this way (to consolidate more control over how we “think” and what we believe)

Re: GPT-5

#912

Ok this[0] sounds very, uh bold to me? Surely this is going to break a ton of workflows etc seemingly with nearly no notice? I'm assuming 'launches' equates with 'fully rolls out' or something but it's not that clear to me. When GPT-5 launches, several older models will be retired, including: - GPT-4o - GPT-4.1 - GPT-4.5 - GPT-4.1-mini - o4-mini - o4-mini-high - o3 - o3-pro If you open a conversation that used one of…

Yeah I was surprised how fast they rugged 4. I guess they want to concentrate their hardware on 5.

If it costs the same compute to run it then there is no point running worse models

Re: GPT-5

#913

Going by the system card at: https://openai.com/index/gpt-5-system-card/ > GPT‑5 is a unified system . . . OK > . . . with a smart and fast model that answers most questions, a deeper reasoning model for harder problems, and a real-time router that quickly decides which model to use based on conversation type, complexity, tool needs, and explicit intent (for example, if you say “think hard about this” in the prompt).…

>This looks like they're not training the single big model but instead have gone off to develop special sub models and attempt to gloss over them with yet another model. That's what you resort to only when doing the end-to-end training has become too expensive for you. The corollary to the bitter lesson strikes again: any hand crafted system will out perform any general system for the same budget by a wide margin.

That is, at best, wishful thinking.

In practice the whole point is the opposite is the case, which is why this direction by OpenAI is a suspicious indicator.

Re: GPT-5

#914

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

AGI will more probably come from google deepmind with a genie model that looks like the matrix moves already

Re: GPT-5

#915

I'm not really convinced, the benchmark blunder was really strange but the demos were quite underwhelming, and it appears this was reflected by a huge market correction in the betting markets as to who will have the best AI by end of the year. What excites me now is that Gemini 3.0 or some answer from Google is coming soon and that will be the one I will actually end up using. It seems like the last mover in the LLM…

Polymarket betters are not impressed. Based upon the market odds, OpenAI had a 35% chance to have the best model (at year end), but those odds have dropped to 18% today. (I'm mostly making this comment to document what happened for the history books.) https://polymarket.com/event/which-company-has-best-ai-model...

You don't actually hold polymarket odds with any significant weighting on actual outcomes do you?

Re: GPT-5

#916

Ok this[0] sounds very, uh bold to me? Surely this is going to break a ton of workflows etc seemingly with nearly no notice? I'm assuming 'launches' equates with 'fully rolls out' or something but it's not that clear to me. When GPT-5 launches, several older models will be retired, including: - GPT-4o - GPT-4.1 - GPT-4.5 - GPT-4.1-mini - o4-mini - o4-mini-high - o3 - o3-pro If you open a conversation that used one of…

> For Free and Plus users, these changes take effect immediately. Pro, Team, and Enterprise users will also see the changes at launch but will have access to older models through legacy model settings. So only for free/plus users (for now). I do wonder how long they will take to deprecate these models via API though...

I'm not worried about when they will deprecate them but I am worried about when they will be removed

3.5 Turbo has been deprecated for a long time but is still running

Re: GPT-5

#917

Watching the livestream now, the improvement over their current models on the benchmarks is very small. I know they seemed to be trying to temper our expectations leading up to this, but this is much less improvement than I was expecting

Law of diminishing returns. We’re talking about less than a 10% performance gain, for a shitload of data, time, and money investment.

I'm not sure what "10% performance gain" is supposed to mean here; but moving from "It does a decent job 95% of the time but screws it up 5%" to "It does a decent job 98% of the time and screws it up 2%" to "It does a decent job 99.5% of the time and only screws it up 0.5%" are major qualitative improvements.

Re: GPT-5

#920

Ok this[0] sounds very, uh bold to me? Surely this is going to break a ton of workflows etc seemingly with nearly no notice? I'm assuming 'launches' equates with 'fully rolls out' or something but it's not that clear to me. When GPT-5 launches, several older models will be retired, including: - GPT-4o - GPT-4.1 - GPT-4.5 - GPT-4.1-mini - o4-mini - o4-mini-high - o3 - o3-pro If you open a conversation that used one of…

> For Free and Plus users, these changes take effect immediately. Pro, Team, and Enterprise users will also see the changes at launch but will have access to older models through legacy model settings. So only for free/plus users (for now). I do wonder how long they will take to deprecate these models via API though...

So they confirmed what we've all been speculating: this is a cost saving update

Smaller base models + more RL. Technically better at the verticals that are making money, but worse on subjective preference.

They'll probably try to prompt engineer back in some of the "vibes", hence the personalities. But also maybe they decided people spending $20 a month to hammer 4o all day as a friend (no judgement, really) are ok to tick off for now... and judging by Reddit, they are very ticked off.

Post reply on HN