GPT-5
261–270 of 1001 posts
Re: GPT-5
#262The silent victory here is this seems like it is being built to be faster and cheaper than o3 while presenting a reasonable jump, which is an important jump in scaling law On the other hand if it's just getting bigger and slower it's not a good sign for LLMs
Re: GPT-5
#263Re: GPT-5
#264> With GPT-5 we will be deprecating all of our prior models Wow, they actually did it
Re: GPT-5
#265Re: GPT-5
#26674.9 SWEBench. This increases the SOTA by a whole .4%. Although the pricing is great, it doesn't seem like OpenAI found a giant breakthrough yet like o1 or Claude 3.5 Sonnet
Re: GPT-5
#267In terms of raw prose quality, I'm not convinced GPT-5 sounds "less like AI" or "more like a friend". Just count the number of em-dashes. It's become something of a LLM shibboleth.
They're all working on subjective improvements, but for example, none of them would develop and deploy a sampler that makes models 50% worse at coding but 50% less likely to use purple prose.
(And unlike the early days where better coding meant better everything, more of the gains are coming from very specific post-training that transfers less, and even harms performance there)
Re: GPT-5
#268Re: GPT-5
#269Seems LLMs really hit the wall.
A whole 8 months ago.
Re: GPT-5
#270Someone at OpenAI screwed up the SWE-bench graph. o3 and GPT-4o bars are same height, but with different values.
It feels a bit intentional