Live data from Hacker News

GPT-5

openai.com

831–840 of 1001 posts

Re: GPT-5

#831
post #750
post #281

> 400,000 context window > 128,000 max output tokens > Input $1.25 > Output $10.00 Source: https://platform.openai.com/docs/models/gpt-5 If this performs well in independent needle-in-haystack and adherence evaluations, this pricing with this context window alone would make GPT-5 extremely competitive with Gemini 2.5 Pro and Claude Opus 4.1, even if the output isn't a significant improvement over o3. If the output qu…

Being on-par with competitors is somehow a "massive leap" for OpenAI now? How far have they fallen...

Are you kidding? If GPT 5 is really on par with Opus 4.1, it means now OpenAI is offering the same product but 10 times cheaper. In any other industry it's not just a massive leap. It's "all competitors are out of market in a few months if they can't release something similar."

Re: GPT-5

#832

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

It's also worth considering that past some threshold, it may be very difficult for us as users to discern which model is better. I don't think thats what's going on here, but we should be ready for it. For example, if you are an ELO 1000 chess player would you yourself be able to tell if Magnus Carlson or another grandmaster were better by playing them individually? To the extent that our AGI/SI metrics are based on…

That's a great point. Thanks.

Re: GPT-5

#833

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

Is it?

Nothing we have is anywhere near AGI and as models age others can copy them.

I personally think we are closing the end of improvement for LLMs with current methods. We have consumed all of the readily available data already, so there is no more good quality training material left. We either need new novel approaches or hope that if enough compute is thrown at training actual intelligence will spontaneously emerge.

Re: GPT-5

#834

I am thoroughly unimpressed by GPT-5. It still can't compose iambic trimeters in ancient Greek with a proper penthemimeral cæsura, and it insists on providing totally incorrect scansion of the flawed lines it does compose. I corrected its metrical sins twice, which sent it into "thinking" mode until it finally returned a "Reasoning failed" error. There is no intelligence here: it's still just giving plausible output.…

I can't tell whether you're serious or not. Your criterion for an "impressive" AI tool is that it be able to write and scan poetry in ancient Greek?

Re: GPT-5

#835
post #198

Watching the livestream now, the improvement over their current models on the benchmarks is very small. I know they seemed to be trying to temper our expectations leading up to this, but this is much less improvement than I was expecting

GPT-5 is #1 on WebDev Arena with +75 pts over Gemini 2.5 Pro and +100 pts over Claude Opus 4: https://lmarena.ai/leaderboard

That eval hasn't been relevant for a while now. Performance there just doesn't seem to correlate well with real-world performance.

Re: GPT-5

#836
post #344

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

Companies are collections of people, and these companies keep losing key developers to the others, I think this is why the clusters happen. OpenAI is now resorting to giving million dollar bonuses to every employee just to try to keep them long term.

that kid at meta negotiated 250m

Re: GPT-5

#837

ChatGPT5 in this demo: > For an airplane wing (airfoil), the top surface is curved and the bottom is flatter. When the wing moves forward: > * Air over the top has to travel farther in the same amount of time -> it moves faster -> pressure on the top decreases. > * Air underneath moves slower -> pressure underneath is higher > * The presure difference creates an upward force - lift Isn't that explanation of why wings…

They did not ask how wings work. They asked for the bernoulli effect, that's a different question.

Re: GPT-5

#839
post #770

So far GPT-5 has not been able to pass my personal "Turing test" which has been unsuccessful for the past several years starting through various versions of Dall-e up to the latest model. I want it to create an image of Santa Claus pulling the sleigh with a reindeer in the sleigh holding the reins, driving the sleigh. No matter how I modify the prompt it is still unable to create this image that my daughter requested…

Is GPT-5 not just routing this request to a 4o/other tool call?

Re: GPT-5

#840
Hypothesis: to the average user this will feel like a much greater jump in capability then to the average HNer, because most users were not using the model selector. So it'll be more successful than the benchmarks suggest.
Post reply on HN