Live data from Hacker News

GPT-5

openai.com

331–340 of 1001 posts

Re: GPT-5

#331
It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie they can all basically solve moderately challenging math and coding problems).

As a user, it feels like the race has never been as close as it is now. Perhaps dumb to extrapolate, but it makes me lean more skeptical about the hard take-off / winner-take-all mental model that has been pushed.

Would be curious to hear the take of a researcher at one of these firms - do you expect the AI offerings across competitors to become more competitive and clustered over the next few years, or less so?

Re: GPT-5

#332
What did Ilya see? (or rather what could he no longer bear to see?)

> Academics distorting graphs to make their benchmarks appear more impressive

> lavish 1.5 million dollar bonuses for everyone at the company

> Releasing an open source model that doesn't even use latent multi head attention in a open source AI world led by Chinese labs

> Constantly overhyping models as scary and dangerous to buy time to lobby against competitors and delay product launches

> Failing to match that hype as AGI is not yet here

Re: GPT-5

#333

ChatGPT5 in this demo: > For an airplane wing (airfoil), the top surface is curved and the bottom is flatter. When the wing moves forward: > * Air over the top has to travel farther in the same amount of time -> it moves faster -> pressure on the top decreases. > * Air underneath moves slower -> pressure underneath is higher > * The presure difference creates an upward force - lift Isn't that explanation of why wings…

Nobody explains it as well as Bartosz: https://ciechanow.ski/airfoil/

Re: GPT-5

#334

What's going on with this plot's y-axis? https://bsky.app/profile/tylermw.com/post/3lvtac5hues2n

It makes it look like the presentation is rushed or made last minute. Really bad to see this as the first plot in the whole presentation. Also, I would have loved to see comparisons with Opus 4.1. Edit: Opus 4.1 scores 74.5% ( https://www.anthropic.com/news/claude-opus-4-1 ). This makes it sound like Anthropic released the upgrade to still be the leader on this important benchmark.

> like the presentation is rushed or made last minute

Or written by GPT-5?

Re: GPT-5

#335

ChatGPT5 in this demo: > For an airplane wing (airfoil), the top surface is curved and the bottom is flatter. When the wing moves forward: > * Air over the top has to travel farther in the same amount of time -> it moves faster -> pressure on the top decreases. > * Air underneath moves slower -> pressure underneath is higher > * The presure difference creates an upward force - lift Isn't that explanation of why wings…

As a complete aside I’ve always hated that explanation where air moves up and over a bump, the lines get closer together and then the explanation is the pressure lowers at that point. Also the idea that the lines of air look the same before and after and yet somehow the wing should have moved up.

Re: GPT-5

#336
They will retire lots of models: GPT-4o, GPT-4.1, GPT-4.5, GPT-4.1-mini, o4-mini, o4-mini-high, o3, o3-pro.

https://help.openai.com/en/articles/6825453-chatgpt-release-...

"If you open a conversation that used one of these models, ChatGPT will automatically switch it to the closest GPT-5 equivalent."

- 4o, 4.1, 4.5, 4.1-mini, o4-mini, or o4-mini-high => GPT-5

- o3 => GPT-5-Thinking

- o3-Pro => GPT-5-Pro

Re: GPT-5

#338

The marketing copy and the current livestream appear tautological: "it's better because it's better." Not much explanation yet why GPT-5 warrants a major version bump. As usual, the model (and potentially OpenAI as a whole) will depend on output vibe checks.

> Not much explanation yet why GPT-5 warrants a major version bump Exactly. Too many videos - too little real data / benchmarks on the page. Will wait for vibe check from simonw and others

> Will wait for vibe check from simonw

https://openai.com/gpt-5/?video=1108156668

2:40 "I do like how the pelican's feet are on the pedals." "That's a rare detail that most of the other models I've tried this on have missed."

4:12 "The bicycle was flawless."

5:30 Re generating documentation: "It nailed it. It gave me the exact information I needed. It gave me full architectural overview. It was clearly very good at consuming a quarter million tokens of rust." "My trust issues are beginning to fall away"

Edit: ohh he has blog post now: https://news.ycombinator.com/item?id=44828264

Re: GPT-5

#339

ChatGPT5 in this demo: > For an airplane wing (airfoil), the top surface is curved and the bottom is flatter. When the wing moves forward: > * Air over the top has to travel farther in the same amount of time -> it moves faster -> pressure on the top decreases. > * Air underneath moves slower -> pressure underneath is higher > * The presure difference creates an upward force - lift Isn't that explanation of why wings…

The hallmark of an LLM response: plausible sounding, but if you dig deeper, incorrect

Re: GPT-5

#340
post #74

What's going on with their SWE bench graph?[0] GPT-5 non-thinking is labeled 52.8% accuracy, but o3 is shown as a much shorter bar, yet it's labeled 69.1%. And 4o is an identical bar to o3, but it's labeled 30.8%... [0] https://i.postimg.cc/DzkZZLry/y-axis.png

Probably generated by an LLM
Post reply on HN