Live data from Hacker News

GPT-5

openai.com

91–100 of 1001 posts

Re: GPT-5

#91
post #45

Does this mean AGI is cancelled? 2027 hard takeoff was just sci-fi?

At this point the prediction for SWE bench (85% by end of this month) is not materializing. We're actually quite far away.

Re: GPT-5

#92
It says out now in chatgpt. Did anyone yet hit the usage limits to report back how many messages are possible?

Re: GPT-5

#93
post #74

What's going on with their SWE bench graph?[0] GPT-5 non-thinking is labeled 52.8% accuracy, but o3 is shown as a much shorter bar, yet it's labeled 69.1%. And 4o is an identical bar to o3, but it's labeled 30.8%... [0] https://i.postimg.cc/DzkZZLry/y-axis.png

The barplot is wrong, the numbers are correct. Looks like they had a dummy plot and never updated it, only the numbers to prevent leaking?

Screenshot of the blog plot: https://imgur.com/a/HAxIIdC

Re: GPT-5

#94

Wait, isn't the Bernoulli effect thing they're demoing now wrong? I thought that was a "common misconception" and wings don't really work by the "longer path" that air takes over the top, and that it was more about angle of attack (which is why planes can fly upside down). It seems like it's actually an ideal "trick" question for an LLM actually, since so much content has been written about it incorrectly. I thought…

That's what I thought. Aeroplanes don't fly because of the Bernoulli effect:

https://physics.stackexchange.com/questions/290/what-really-...

Apparently. Not that I know either way.

Re: GPT-5

#95
74.9 SWEBench. This increases the SOTA by a whole .4%. Although the pricing is great, it doesn't seem like OpenAI found a giant breakthrough yet like o1 or Claude 3.5 Sonnet

Re: GPT-5

#96
post #74

What's going on with their SWE bench graph?[0] GPT-5 non-thinking is labeled 52.8% accuracy, but o3 is shown as a much shorter bar, yet it's labeled 69.1%. And 4o is an identical bar to o3, but it's labeled 30.8%... [0] https://i.postimg.cc/DzkZZLry/y-axis.png

GPT-5 generated the chart

Re: GPT-5

#97
post #92

It says out now in chatgpt. Did anyone yet hit the usage limits to report back how many messages are possible?

I don't see it in my model picker yet.

Re: GPT-5

#98
post #28

# GPT5 all official links Livestream link: https://www.youtube.com/live/0Uu_VJeVVfo Research blog post: https://openai.com/index/introducing-gpt-5/ Developer blog post: https://openai.com/index/introducing-gpt-5-for-developers API Docs: https://platform.openai.com/docs/guides/latest-model Note the free form function calling documentation: https://platform.openai.com/docs/guides/function-calling#con... GPT5 prompting…

our hands on review: https://www.latent.space/p/gpt-5-review

basically in my testing really felt that gpt5 was "using tools to think" rather than just "using tools". it gets very powerful when coding long horizon tasks (a separate post i'm publishing later).

to give one substantive example, in my developer beta (they will release the video in a bit) i put it to a task that claude code had been stuck on for the last week - same prompts - and it just added logging to instrument some of the failures that we were seeing and - from the logs that it added and asked me to rerun - figured out the solve.

Re: GPT-5

#100

Note it's not available to everyone yet: > GPT-5 Rollout > We are gradually rolling out GPT-5 to ensure stability during launch. Some users may not yet see GPT-5 in their account as we increase availability in stages.

But available from today to free tier. Yay.
Post reply on HN