Live data from Hacker News

GPT-5

openai.com

221–230 of 1001 posts

Re: GPT-5

#221

ChatGPT5 in this demo: > For an airplane wing (airfoil), the top surface is curved and the bottom is flatter. When the wing moves forward: > * Air over the top has to travel farther in the same amount of time -> it moves faster -> pressure on the top decreases. > * Air underneath moves slower -> pressure underneath is higher > * The presure difference creates an upward force - lift Isn't that explanation of why wings…

[deleted]

Re: GPT-5

#223

Seems LLMs really hit the wall.

Before last year we didn't have reasoning. It came with QuietSTaR, then we got it in the form of O1 and then it became practical with DeepSeek's paper in January.

So we're only about a year since the last big breakthrough.

I think we got a second big breakthrough with Google's results on the IMO problems.

For this reason I think we're very far from hitting a wall. Maybe 'LLM parameter scaling is hitting a wall'. That might be true.

Re: GPT-5

#224
The ultimate test I’ve found so far is to create OpenSCAD models with the LLM. They really struggle with the mapping 3D space objects. Curious to see how GPT-5 is performs here.

Re: GPT-5

#227

But can it say “I don’t know” if ya know, it doesn’t

I agree with the sentiment, but the problem with this question is that LLMs don't "know" *anything*, and they don't actually "know" how to answer a question like this.

It's just statistical text generation. There is *no actual knowledge*.

Re: GPT-5

#228
post #198

Watching the livestream now, the improvement over their current models on the benchmarks is very small. I know they seemed to be trying to temper our expectations leading up to this, but this is much less improvement than I was expecting

GPT-5 is #1 on WebDev Arena with +75 pts over Gemini 2.5 Pro and +100 pts over Claude Opus 4: https://lmarena.ai/leaderboard

This same leaderboard lists a bunch of models, including 4o, beating out Opus 4, which seems off.

Re: GPT-5

#229
post #137

The marketing copy and the current livestream appear tautological: "it's better because it's better." Not much explanation yet why GPT-5 warrants a major version bump. As usual, the model (and potentially OpenAI as a whole) will depend on output vibe checks.

We’re at the audiophile stage of LLMs where people are talking about the improved soundstage, tonality, reduced sibilance etc

It’s always been this way with LLMs.

Re: GPT-5

#230
Someone at OpenAI screwed up the SWE-bench graph. o3 and GPT-4o bars are same height, but with different values.
Post reply on HN