Live data from Hacker News

GPT-5

openai.com

51–60 of 1001 posts

Re: GPT-5

#51
The silent victory here is this seems like it is being built to be faster and cheaper than o3 while presenting a reasonable jump, which is an important jump in scaling law

On the other hand if it's just getting bigger and slower it's not a good sign for LLMs

Re: GPT-5

#52
post #28

# GPT5 all official links Livestream link: https://www.youtube.com/live/0Uu_VJeVVfo Research blog post: https://openai.com/index/introducing-gpt-5/ Developer blog post: https://openai.com/index/introducing-gpt-5-for-developers API Docs: https://platform.openai.com/docs/guides/latest-model Note the free form function calling documentation: https://platform.openai.com/docs/guides/function-calling#con... GPT5 prompting…

Aaand hugged to death.

edit:

livestream here: https://www.youtube.com/live/0Uu_VJeVVfo

Re: GPT-5

#53
So the benchmark graphs they have shown so far in the stream appears to show that GPT-5 is WORSE than other models unless you use thinking?

Re: GPT-5

#56
They claim it thinks the "perfect amount" but there is no perfect amount. It all depends on willingness to pay, latency tolerance, etc.

Re: GPT-5

#57
post #43

The Polyglot aider improvement over o3 is imperceptible, not great.

SWE-Bench is also not stellar. "It's important to remember" that:

- they are only evals

- this is mostly positioned as a general consumer product, they might have better stuff for us nerds in hand.

Re: GPT-5

#58

Seems LLMs really hit the wall.

It's seemed that way for the last year. The only real improvements have been in the chat apps themselves (internet access, function calling). Until AI gets past the pre-training problem, it'll stagnate.

Re: GPT-5

#59
post #47

I wish they posted detailed metrics and benchmarks with such a "big" (loud) update.

The current livestream listed the benchmarks (curiously comparing it only to previous GPT models and not competitors)

Re: GPT-5

#60
The dev blog makes it sound like they’re aiming more for “AI teammate” than just another upgrade. That said, it’s hard to tell how much of this is real improvement vs better packaging. Benchmarks are cherry-picked as usual, and there’s not much comparison to other models. Curious to hear how it performs in actual workflows.
Post reply on HN