Live data from Hacker News

GPT-5

openai.com

981–990 of 1001 posts

Re: GPT-5

#981

I'm not really convinced, the benchmark blunder was really strange but the demos were quite underwhelming, and it appears this was reflected by a huge market correction in the betting markets as to who will have the best AI by end of the year. What excites me now is that Gemini 3.0 or some answer from Google is coming soon and that will be the one I will actually end up using. It seems like the last mover in the LLM…

Polymarket betters are not impressed. Based upon the market odds, OpenAI had a 35% chance to have the best model (at year end), but those odds have dropped to 18% today. (I'm mostly making this comment to document what happened for the history books.) https://polymarket.com/event/which-company-has-best-ai-model...

> Polymarket betters are not impressed. Based upon the market odds, OpenAI had a 35% chance to have the best model (at year end)

who will decide the winner to resolve bets?

Re: GPT-5

#982

Ok this[0] sounds very, uh bold to me? Surely this is going to break a ton of workflows etc seemingly with nearly no notice? I'm assuming 'launches' equates with 'fully rolls out' or something but it's not that clear to me. When GPT-5 launches, several older models will be retired, including: - GPT-4o - GPT-4.1 - GPT-4.5 - GPT-4.1-mini - o4-mini - o4-mini-high - o3 - o3-pro If you open a conversation that used one of…

I don't think they build the ChatGPT subscription interface around workflows. They leave the APIs to cover either custom or pre-built workflows pinned against a specific model. At the same time I wouldn't be surprised if they do end up keeping a couple of the cheaper/smaller older models though, the cost would be low and reduce a lot of the churn friction.

I'm not saying I'd do it that way myself, but it explains why they don't see it as too bold.

Re: GPT-5

#983

Is anyone else having problems with factual correctness? I had a number of 4o and o3 conversations going and those models were factually correct about a number of different subjects. Asking GPT-5 about the same things results in wrong answers even though its training data is newer. And it won't look things up to correct itself unless I manually switch to the thinking variant. This is worse. I cancelled my subscriptio…

Example?

Re: GPT-5

#984

GPT-5 knowledge cutoff: Sep 30, 2024 (10 months before release). Compare that to Gemini 2.5 Pro knowledge cutoff: Jan 2025 (3 months before release) Claude Opus 4.1: knowledge cutoff: Mar 2025 (4 months before release) https://platform.openai.com/docs/models/compare https://deepmind.google/models/gemini/pro/ https://docs.anthropic.com/en/docs/about-claude/models/overv...

Does the knowledge cut off date still matter all that much since all these models can do real time searches and RAG?

Re: GPT-5

#986

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

The inflection point is recursive self-improvement. Once an AI achieves that, and I mean really achieves it - where it can start developing and deploying novel solutions to deep problems that currently bottleneck its own capabilities - that's where one would suddenly leap out in front of the pack and then begin extending its lead. Nobody's there yet though, so their performance is clustering around an asymptotic limit of what LLMs are capable of.

Re: GPT-5

#987

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

If AGI is ever achieved, it would open the door to recursive self improvement that would presumably rapidly exceed human capability across any and all fields, including AI development. So the AI would be improving itself while simultaneously also making revolutionary breakthroughs in essentially all fields. And, for at least a while, it would also presumably be doing so at an exponentially increasing rate. But I thin…

> But I think we're not even on the path to creating AGI.

It seems like the LLM model will be component of an eventual AGI, it's voice per se, but not its mind. The mind still requires another innovation or breakthrough we haven't seen yet.

Re: GPT-5

#988

Ok this[0] sounds very, uh bold to me? Surely this is going to break a ton of workflows etc seemingly with nearly no notice? I'm assuming 'launches' equates with 'fully rolls out' or something but it's not that clear to me. When GPT-5 launches, several older models will be retired, including: - GPT-4o - GPT-4.1 - GPT-4.5 - GPT-4.1-mini - o4-mini - o4-mini-high - o3 - o3-pro If you open a conversation that used one of…

This behavior should be an early warning sign of future potential enshitification and a reason to consider open weight models you can host elsewhere. If you are building on models that could disappear tomorrow when a company needs to juice the launch of a new model (or increase prices), you are introducing avoidable risk.

This was my read as well.

Doesn't matter at all if the newer model is earth-shatteringly good (and this one doesn't seem to be): If I can't reliably access the models I've built my tooling on top of... I'm very unhappy.

If this note is just intended for the GUI chat interface they provide - Fine. I don't love it, but I get it.

But if the older models start disappearing from the paid API surfaces (ex - I can no longer get to a precise snapshot through something like "gpt-4o-2024-08-06" or "gpt-3.5-turbo-1106") then this is a great reason to abandon OpenAI entirely as a platform.

Re: GPT-5

#989

They vibe coded the update. "Your organization must be verified to use the model `gpt-5`. Please go to: https://platform.openai.com/settings/organization/general and click on Verify Organization. If you just verified, it can take up to 15 minutes for access to propagate." And every way I click through this I end in an infinity loop on the site...

So, first it did not work because of API changes. Then I got the problem with the loop. And then it did not work either cause it requires withpersona.

Re: GPT-5

#990
post #470

I did a little test that I like to do with new models: "I have rectangular space of dimensions 30x30x90mm. Would 36x14x60mm battery fit in it, show in drawing proof". GPT5 failed spectacularly.

I tried it again today out of curiosity. OpenAI said there was some routing bug on launch and requests were going to the cheaper model.

Today it seems pretty good. Not perfect, but not a spectacular failure.

https://chatgpt.com/s/t_68966fcf457c8191811968b9a6a2e81e

Post reply on HN