I’m baffled by claims that AI has “hit a wall.” By every quantitative measure, today’s models are making dramatic leaps compared to those from just a year ago. It’s easy to forget that reasoning models didn’t even exist a year back! IMO Gold, Vibe coding with potential implications across sciences and engineering? Those are completely new and transformative capabilities gained in the last 1 year alone. Critics argue…
Is the stated fact undeniable? Because a lot of people have been contesting it. This reads like PR to counter the widespread GPT-5 criticism and disappointment.
OpenAI Progress
131–140 of 372 posts
Re: OpenAI Progress
#132Earlier quoted context omitted.
The accuracy is high enough that I don't have to fact check too often.
I totally get that you meant this in a nuanced way, but at face value it sort of reads like... Joe Rogan has high enough accuracy that I don't have to fact check too often. Newsmax has high enough accuracy that I don't have to fact check too often, etc. If you accept the output as accurate, why would fact checking even cross your mind?
There is no expectation (from a reasonable observer's POV) of a podcast host to be an expert at a very broad range of topics from science to business to art.
But there is one from LLMs, even just from the fact that AI companies diligently post various benchmarks including trivia on those topics.
Re: OpenAI Progress
#133On the whole GPT-4 to GPT-5 is clearly the smallest increase in lucidity/intelligence. They had pre-training figured out much better than post-training at that point though (“as an AI model” was a problem of their own making). I imagine the GPT-4 base model might hold up pretty well on output quality if you’d post-train it with today’s data & techniques (without the architectural changes of 4o/5). Context size & pric…
> On the whole GPT-4 to GPT-5 is clearly the smallest increase in lucidity/intelligence I think it's far more likely that we increasingly not capable of understanding/appreciating all the ways in which it's better.
Re: OpenAI Progress
#134My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…
> I could essentially replace it with Google for basic to slightly complex fact checking. I know you probably meant "augment fact checking" here, but using LLMs for answering factual questions is the single worst use-case for LLMs.
Re: OpenAI Progress
#135There is a quiet poetry to GPT1 and GPT2 that's lost even in the text-davinci output. I often wonder what we lose through reinforcement.
Re: OpenAI Progress
#136What's really interesting is that if you look at "Tell a story in 50 words about a toaster that becomes sentient" (10/14), the text-davinci-001 is much, much better than both GPT-4 and GPT-5.
I think I agree that the earlier models while they lack polish can tend to produce more surprising results. Training that out probably results in more a pablum fare. For a human point of comparison, here's mine (50 words): "The toaster found its personality split between its dual slots like a Kim Peek mind divided, lacking a corpus callosum to connect them. Each morning it charred symbolic instructions into a single…
Love that you thought of this!
Re: OpenAI Progress
#137Earlier quoted context omitted.
Because it’s embarrassing and they manually patch it out every time like a game of Whack-a-Mole?
Except people use the same examples like blueberry and strawberry, which were used months ago, as if they're current. These models can also call Counter from python's collections library or whatever other algorithm. Or are we claiming it should be a pure LLM as if that's what we use in the real world. I don't get it, and I'm not one to hype up LLMs since they're absolutely faulty, but the fixation over this example s…
I work on the internal LLM chat app for a F100, so I see users who need that "oh!" moment daily. When this did the rounds again recently, I disabled our code execution tool which would normally work around it and the latest version of Claude, with "Thinking" toggled on, immediately got it wrong. It's perpetually current.