Live data from Hacker News

GPT-5

openai.com

141–150 of 1001 posts

Re: GPT-5

#141

Sam Altman, in the summer update video: > "[GPT-5] can write an entire computer program from scratch, to help you with whatever you'd like. And we think this idea of software on demand is going to be one of the defining characteristics of the GPT-5 era."

Cannot believe how it could stand up to that high expectation.

But then again, all of this is a hype machine cranked up till the next one needs cranking.

Re: GPT-5

#142
My conspiracy theory is that the introductory footage of Sam in this and the Jony Ive video is AI generated

Re: GPT-5

#143
In terms of raw prose quality, I'm not convinced GPT-5 sounds "less like AI" or "more like a friend". Just count the number of em-dashes. It's become something of a LLM shibboleth.

Re: GPT-5

#146

Wait, isn't the Bernoulli effect thing they're demoing now wrong? I thought that was a "common misconception" and wings don't really work by the "longer path" that air takes over the top, and that it was more about angle of attack (which is why planes can fly upside down). It seems like it's actually an ideal "trick" question for an LLM actually, since so much content has been written about it incorrectly. I thought…

Yeah, they sure clicked away from it very fast and kept adjusting the scrollbars. It was confusing what it was trying to display. Furthermore, the prompt contained "Canvas" and "SVG" while as someone with webdev experience these are certainly familiar concepts, i wouldn't consider those in the "casual lexicon" for a random user trying to help a middle schooler with homework. I'm not impressed...

IMO Claude 3.7 could have done a similar / better job with that a year ago.

Re: GPT-5

#147
I dont know if there is a faster way to get me riled up: say 'try it' (me a Pro member) and then not getting it because I am logged in. Got opus 4.1 when it appeared. Not sure what is happening here but I am out.

Re: GPT-5

#150
post #74

What's going on with their SWE bench graph?[0] GPT-5 non-thinking is labeled 52.8% accuracy, but o3 is shown as a much shorter bar, yet it's labeled 69.1%. And 4o is an identical bar to o3, but it's labeled 30.8%... [0] https://i.postimg.cc/DzkZZLry/y-axis.png

They vibecharted
Post reply on HN