Live data from Hacker News

GPT-5

openai.com

81–90 of 1001 posts

Re: GPT-5

#81
post #2

For day to day coding, I've found Anthropic to be killing it with Sonnet 3.7 and now Sonnet 4, and Claude Code feeling like it has even bigger advantages over when it's used in Cursor (And I can't explain why). I don't even try to use the OpenAI models because it's felt like night and day. Hopefully GPT-5 helps them catch up. Although I'm sure there are 100 people that have their own personal "hopefully GPT-5 fixes m…

Killing it - at what type of coding task? What "bigger advantages" specifically? What is night and day?

Re: GPT-5

#83
post #36

It's very interesting how memetic the language around different models is. Elon seems to have coined "PhD level intelligence in all topics" and now Sam repeated it in his presentation. Despite it not having an actual meaning. I think OpenAI will coin they've achieved AGI first (as they have incentives to based on the rumored contract with MSFT), and then everyone will claim we've achieved it.

Elon did not coin this, Kurzweil has been using this coinage for a lot longer.

Re: GPT-5

#84
The presentation asks for a moving svg to illustrate Bernoulli, that's suspiciously close to a Pelican.

Re: GPT-5

#85
post #12

> comments turned off yikes - the poor executive leadership’s fragile egos cannot take the criticism.

What's up with their very first eval? The SWE bars and numbers don't line up.

Re: GPT-5

#86
Wait, isn't the Bernoulli effect thing they're demoing now wrong? I thought that was a "common misconception" and wings don't really work by the "longer path" that air takes over the top, and that it was more about angle of attack (which is why planes can fly upside down).

It seems like it's actually an ideal "trick" question for an LLM actually, since so much content has been written about it incorrectly. I thought at first they were going to demo this to show that it knew better, but it seems like it's just regurgitating the same misleading stuff. So, not a good look.

Re: GPT-5

#90
post #74

What's going on with their SWE bench graph?[0] GPT-5 non-thinking is labeled 52.8% accuracy, but o3 is shown as a much shorter bar, yet it's labeled 69.1%. And 4o is an identical bar to o3, but it's labeled 30.8%... [0] https://i.postimg.cc/DzkZZLry/y-axis.png

also wondering this. had to pause the livestream to make sure i wasnt crazy. definitely eyebrow raising
Post reply on HN