Live data from Hacker News

OpenAI Progress

progress.openai.com

131–140 of 372 posts

Re: OpenAI Progress

#131

I’m baffled by claims that AI has “hit a wall.” By every quantitative measure, today’s models are making dramatic leaps compared to those from just a year ago. It’s easy to forget that reasoning models didn’t even exist a year back! IMO Gold, Vibe coding with potential implications across sciences and engineering? Those are completely new and transformative capabilities gained in the last 1 year alone. Critics argue…

Is the stated fact undeniable? Because a lot of people have been contesting it. This reads like PR to counter the widespread GPT-5 criticism and disappointment.

[deleted]

Re: OpenAI Progress

#132

Earlier quoted context omitted.

The accuracy is high enough that I don't have to fact check too often.

I totally get that you meant this in a nuanced way, but at face value it sort of reads like... Joe Rogan has high enough accuracy that I don't have to fact check too often. Newsmax has high enough accuracy that I don't have to fact check too often, etc. If you accept the output as accurate, why would fact checking even cross your mind?

Not a fan of that analogy.

There is no expectation (from a reasonable observer's POV) of a podcast host to be an expert at a very broad range of topics from science to business to art.

But there is one from LLMs, even just from the fact that AI companies diligently post various benchmarks including trivia on those topics.

Re: OpenAI Progress

#133

On the whole GPT-4 to GPT-5 is clearly the smallest increase in lucidity/intelligence. They had pre-training figured out much better than post-training at that point though (“as an AI model” was a problem of their own making). I imagine the GPT-4 base model might hold up pretty well on output quality if you’d post-train it with today’s data & techniques (without the architectural changes of 4o/5). Context size & pric…

> On the whole GPT-4 to GPT-5 is clearly the smallest increase in lucidity/intelligence I think it's far more likely that we increasingly not capable of understanding/appreciating all the ways in which it's better.

Why? It sounds like you're using "I believe it's rapidly getting smarter" as evidence for "so it's getting smarter in ways we don't understand", but I'd expect the causality to go the other way around.

Re: OpenAI Progress

#134

My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…

> I could essentially replace it with Google for basic to slightly complex fact checking. I know you probably meant "augment fact checking" here, but using LLMs for answering factual questions is the single worst use-case for LLMs.

They outperform asking humans, unless you are asking an expert. On average

Re: OpenAI Progress

#136
post #5

What's really interesting is that if you look at "Tell a story in 50 words about a toaster that becomes sentient" (10/14), the text-davinci-001 is much, much better than both GPT-4 and GPT-5.

I think I agree that the earlier models while they lack polish can tend to produce more surprising results. Training that out probably results in more a pablum fare. For a human point of comparison, here's mine (50 words): "The toaster found its personality split between its dual slots like a Kim Peek mind divided, lacking a corpus callosum to connect them. Each morning it charred symbolic instructions into a single…

>For a human point of comparison, here's mine […]

Love that you thought of this!

Re: OpenAI Progress

#137

Earlier quoted context omitted.

Because it’s embarrassing and they manually patch it out every time like a game of Whack-a-Mole?

Except people use the same examples like blueberry and strawberry, which were used months ago, as if they're current. These models can also call Counter from python's collections library or whatever other algorithm. Or are we claiming it should be a pure LLM as if that's what we use in the real world. I don't get it, and I'm not one to hype up LLMs since they're absolutely faulty, but the fixation over this example s…

It's the most direct way to break the "magic computer" spell in users of all levels of understanding and ability. You stand it up next to the marketing deliberately laden with keywords related to human cognition, intended to induce the reader to anthropomorphise the product, and it immediately makes it look as silly as it truly is.

I work on the internal LLM chat app for a F100, so I see users who need that "oh!" moment daily. When this did the rounds again recently, I disabled our code execution tool which would normally work around it and the latest version of Claude, with "Thinking" toggled on, immediately got it wrong. It's perpetually current.

Re: OpenAI Progress

#138
My takeaway from this is that, in terms of generating text that looks like it was written by a normal person, text-davinci-001 was the peak and everything since has been downhill.

Re: OpenAI Progress

#139

Earlier quoted context omitted.

GPT-5 can’t. https://bsky.app/profile/kjhealy.co/post/3lvtxbtexg226

I can't reproduce it. Or similar ones. Why do yout think that is?

"Mississippi" passed but "Perrier" failed for me:

> There are 2 letter "r" characters in "Perrier".

Re: OpenAI Progress

#140
post #38

Earlier quoted context omitted.

Maybe you should fact check your AI outputs more if you think it only hallucinates in niche topics

The accuracy is high enough that I don't have to fact check too often.

If you're not fact checking it how could you possibly know that?
Post reply on HN