Live data from Hacker News

GPT-5

openai.com

251–260 of 1001 posts

Re: GPT-5

#251
The live stream just has Altman interviewing a lady who was diagnosed 3 different cancers.

GPT4 gave her better response than doctors she said.

Re: GPT-5

#252
post #44

[flagged]

There's no way to directly contact another user other than by replying to a post of theirs and hoping for the best.

If you email us at hn@ycombinator.com and tell us who you want to contact, we might be able to email them and ask if they would be willing to have you contact them. No guarantees though!

Re: GPT-5

#253
post #74

What's going on with their SWE bench graph?[0] GPT-5 non-thinking is labeled 52.8% accuracy, but o3 is shown as a much shorter bar, yet it's labeled 69.1%. And 4o is an identical bar to o3, but it's labeled 30.8%... [0] https://i.postimg.cc/DzkZZLry/y-axis.png

Don't ask questions, just consume product.

Re: GPT-5

#254

ChatGPT5 in this demo: > For an airplane wing (airfoil), the top surface is curved and the bottom is flatter. When the wing moves forward: > * Air over the top has to travel farther in the same amount of time -> it moves faster -> pressure on the top decreases. > * Air underneath moves slower -> pressure underneath is higher > * The presure difference creates an upward force - lift Isn't that explanation of why wings…

Yes, it is completely wrong. If this were a valid explanation, flat-plate airfoils could not generate lift. (They can.)

Source: PhD on aircraft design

Re: GPT-5

#255

Pricing seems good, but the open question is still on tool calling reliability. Input: $1.25 / 1M tokens Cached: $0.125 / 1M tokens Output: $10 / 1M tokens With 74.9% on SWE-bench, this inches out Claude Opus 4.1 at 74.5%, but at a much cheaper cost. For context, Claude Opus 4.1 is $15 / 1M input tokens and $75 / 1M output tokens. > "GPT-5 will scaffold the app, write files, install dependencies as needed, and show a…

I switched immediately because of pricing, input token heavy load, but it doesn't even work. For some reason they completely broke the already amateurish API.

Re: GPT-5

#256

Watching the livestream now, the improvement over their current models on the benchmarks is very small. I know they seemed to be trying to temper our expectations leading up to this, but this is much less improvement than I was expecting

Sam said maybe two years ago that they want to avoid "mic drop" releases, and instead want to stick to incremental steps. This is day one, so there is probably another 10-20% in optimizations that can be squeezed out of it in the coming months.

Then why increment the version number here? This is clearly styled like a "mic drop" release but without the numbers to back it up. It's a really bad look when comparing the crazy jump from GPT3 to GPT4 to this slight improvement with GPT5.

Re: GPT-5

#257
The reduction in hallucinations seems like potentially the biggest upgrade. If it reduces hallucinations by 75% or more over o3 and GPT-4o as the graphs claim, it will be a giant step forward. The inability to trust answers given by AI is the biggest single hurdle to clear for many applications.

Re: GPT-5

#258
post #74

What's going on with their SWE bench graph?[0] GPT-5 non-thinking is labeled 52.8% accuracy, but o3 is shown as a much shorter bar, yet it's labeled 69.1%. And 4o is an identical bar to o3, but it's labeled 30.8%... [0] https://i.postimg.cc/DzkZZLry/y-axis.png

someone copy pasted the 3rd bar to the 2nd

Re: GPT-5

#260

Damn, you guys are toxic. So -- they did not invent AGI yet. Yet, I like what I'm seeing. Major progress on multiple fronts. Hallucination fix is exciting on its own. The React demos were mindblowing.

Yeah, when it becomes cool to be anti AI or anti anything in HN for that matter, the takes start becoming ridiculous, if you just think back a couple of years, or even months ago and where we're now and you can't see it, I guess you're just dead set on dying on that hill.
Post reply on HN