Live data from Hacker News

GPT-5

openai.com

61–70 of 1001 posts

Re: GPT-5

#61
Watching the livestream now, the improvement over their current models on the benchmarks is very small. I know they seemed to be trying to temper our expectations leading up to this, but this is much less improvement than I was expecting

Re: GPT-5

#62
The benchmarks in the stream appears to show that GPT-5 performs WORSE than other models unless you enable thinking?

Re: GPT-5

#63
What surprises me the most is that there is no benchmarks table right at the top. Maybe the improvements are not to call home about?

Re: GPT-5

#65

The marketing copy and the current livestream appear tautological: "it's better because it's better." Not much explanation yet why GPT-5 warrants a major version bump. As usual, the model (and potentially OpenAI as a whole) will depend on output vibe checks.

> Not much explanation yet why GPT-5 warrants a major version bump

Exactly. Too many videos - too little real data / benchmarks on the page. Will wait for vibe check from simonw and others

Re: GPT-5

#66

The marketing copy and the current livestream appear tautological: "it's better because it's better." Not much explanation yet why GPT-5 warrants a major version bump. As usual, the model (and potentially OpenAI as a whole) will depend on output vibe checks.

its >o3 performance at gpt4 price. seems pretty obvious

Re: GPT-5

#67

The marketing copy and the current livestream appear tautological: "it's better because it's better." Not much explanation yet why GPT-5 warrants a major version bump. As usual, the model (and potentially OpenAI as a whole) will depend on output vibe checks.

[deleted]

Re: GPT-5

#68
Note it's not available to everyone yet:

> GPT-5 Rollout

> We are gradually rolling out GPT-5 to ensure stability during launch. Some users may not yet see GPT-5 in their account as we increase availability in stages.

Re: GPT-5

#69
Not live for me in the UK. "Try it in ChatGPT" takes me to the normal page and there's no v5 listed in the dropdown.

Re: GPT-5

#70
SWE-Bench Verified score, with thinking, ties Opus 4.1 without thinking.

AIME scores do not appear too impressive at first glance.

They are downplaying benchmarks heavily in the live stream. This was the lab that has been flexing benchmarks as headline figures since forever.

This is a product-focused update. There is no significant jump in raw intelligence or agentic behavior against SOTA.

Post reply on HN