Live data from Hacker News

GPT-5

openai.com

571–580 of 1001 posts

Re: GPT-5

#571

Watching the livestream now, the improvement over their current models on the benchmarks is very small. I know they seemed to be trying to temper our expectations leading up to this, but this is much less improvement than I was expecting

> I know they seemed to be trying to temper our expectations leading up to this

Before the release of the model Sam Altman tweeted a picture of the Death Star appearing over the horizon of a planet.

Re: GPT-5

#572

Watching the livestream now, the improvement over their current models on the benchmarks is very small. I know they seemed to be trying to temper our expectations leading up to this, but this is much less improvement than I was expecting

Law of diminishing returns.

We’re talking about less than a 10% performance gain, for a shitload of data, time, and money investment.

Re: GPT-5

#573

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

I have been saying this before: S-curves look a lot like exponential curves in the beginning.

Thus, it’s easy to mistake one for the other - at least initially.

Re: GPT-5

#574

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

I think it's very fortunate, because I used to be an AI doomer. I still kinda am, but at least I'm now about 70% convinced that the current technological paradigm is not going to lead us to a short-term AI apocalypse. The fortunate thing is that we managed to invent an AI that is good at _copying us_ instead of being a truly maveric agent, which kinda limits it to the "average human" output. However, I still think th…

It won't lead us to an apocalypse apocalypse, but it may well lead us to an economic crisis.

Re: GPT-5

#577
I ran the below prompt to both Kimi2 and GPT5.

how many rs in cranberry?

-- GPT5's response: The word cranberry has two “r”s. One in cran and one in berry.

Kimi2's response: There are three letter rs in the word "cranberry".

Re: GPT-5

#578
post #336

They will retire lots of models: GPT-4o, GPT-4.1, GPT-4.5, GPT-4.1-mini, o4-mini, o4-mini-high, o3, o3-pro. https://help.openai.com/en/articles/6825453-chatgpt-release-... "If you open a conversation that used one of these models, ChatGPT will automatically switch it to the closest GPT-5 equivalent." - 4o, 4.1, 4.5, 4.1-mini, o4-mini, or o4-mini-high => GPT-5 - o3 => GPT-5-Thinking - o3-Pro => GPT-5-Pro

"GPT-4o, GPT-4.1, GPT-4.5, GPT-4.1-mini, o4-mini, o4-mini-high, o3, o3-pro"

The names of GPT models are just terrible. o3 is better than 4o, maybe?

Re: GPT-5

#579

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

What is the AGI threshold? That the model can manage its own self improvement better than humans can? Then the roles will be reversed -- LLM prompting the meat machines to pave its way.

Re: GPT-5

#580

GPT-5 knowledge cutoff: Sep 30, 2024 (10 months before release). Compare that to Gemini 2.5 Pro knowledge cutoff: Jan 2025 (3 months before release) Claude Opus 4.1: knowledge cutoff: Mar 2025 (4 months before release) https://platform.openai.com/docs/models/compare https://deepmind.google/models/gemini/pro/ https://docs.anthropic.com/en/docs/about-claude/models/overv...

with web search, is knowledge cutoff really relevant anymore? Or is this more of a comment on how long it took them to do post-training?

[deleted]
Post reply on HN