Live data from Hacker News

GPT-5

openai.com

181–190 of 1001 posts

Re: GPT-5

#181

The introduction said to try the following prompt Describe me based on all our chats — make it catchy! It was flattering as all get out, but fairly accurate (IMHO) Mike Warot: The Tinkerer of Tomorrow A hardware hacker with a poet’s soul, Mike blends old-school radio wisdom with cutting-edge curiosity. Whether he's decoding atomic clocks, reinventing FPGA logic with BitGrid, or pondering the electromagnetic vector po…

I genuinely believe you are a kickass person, but that text is full of LLM-isms. Listing things, contrasting or reinforcing prallel sentence structures, it even has the dreaded em-dash.

Here's a suprprisingly enlightening (at least to me) video on how to spot LLM writing:

https://www.youtube.com/watch?v=9Ch4a6ffPZY

Re: GPT-5

#183

These presenters all give off such a “sterile” vibe

They look nervous, messing this presentation up could cost them their high-paying jobs.

Re: GPT-5

#184
ChatGPT5 in this demo:

> For an airplane wing (airfoil), the top surface is curved and the bottom is flatter. When the wing moves forward:

> * Air over the top has to travel farther in the same amount of time -> it moves faster -> pressure on the top decreases.

> * Air underneath moves slower -> pressure underneath is higher

> * The presure difference creates an upward force - lift

Isn't that explanation of why wings work completely wrong? There's nothing that forces the air to cover the top distance in the same time that it covers the bottom distance, and in fact it doesn't. https://www.cam.ac.uk/research/news/how-wings-really-work

Very strange to use a mistake as your first demo, especially while talking about how it's phd level.

Re: GPT-5

#185

Watching the livestream now, the improvement over their current models on the benchmarks is very small. I know they seemed to be trying to temper our expectations leading up to this, but this is much less improvement than I was expecting

It is at least much cheaper and seems faster.

They also announced gpt-5-pro but I haven't seen benchmarks on that yet.

Re: GPT-5

#186

What's going on with this plot's y-axis? https://bsky.app/profile/tylermw.com/post/3lvtac5hues2n

It makes it look like the presentation is rushed or made last minute. Really bad to see this as the first plot in the whole presentation. Also, I would have loved to see comparisons with Opus 4.1.

Edit: Opus 4.1 scores 74.5% (https://www.anthropic.com/news/claude-opus-4-1). This makes it sound like Anthropic released the upgrade to still be the leader on this important benchmark.

Re: GPT-5

#187

SWE-Bench Verified score, with thinking, ties Opus 4.1 without thinking. AIME scores do not appear too impressive at first glance. They are downplaying benchmarks heavily in the live stream. This was the lab that has been flexing benchmarks as headline figures since forever. This is a product-focused update. There is no significant jump in raw intelligence or agentic behavior against SOTA.

they aren't downplaying anything.

Re: GPT-5

#189

Watching the livestream now, the improvement over their current models on the benchmarks is very small. I know they seemed to be trying to temper our expectations leading up to this, but this is much less improvement than I was expecting

I mean that's just the consequence of releasing a new model every couple months. If Open AI stayed mostly silent since the GPT-4 release (like they did for most iterations) and only now released 5 then nobody would be complaining about weak gains in benchmarks.

Well it was their choice to call it GPT 5 and not GPT 4.2.

Re: GPT-5

#190

I know that the number is mostly marketing, but are they forced to call it 5 because of external pressure. This seems more like a GPT 4.x

Aren't all LLMs just vibe-versioned?

I can't even define what a (semantic) major version bump would look like.

Post reply on HN