Live data from Hacker News

OpenAI Progress

progress.openai.com

91–100 of 372 posts

Re: OpenAI Progress

#91

On the whole GPT-4 to GPT-5 is clearly the smallest increase in lucidity/intelligence. They had pre-training figured out much better than post-training at that point though (“as an AI model” was a problem of their own making). I imagine the GPT-4 base model might hold up pretty well on output quality if you’d post-train it with today’s data & techniques (without the architectural changes of 4o/5). Context size & pric…

> On the whole GPT-4 to GPT-5 is clearly the smallest increase in lucidity/intelligence

I think it's far more likely that we increasingly not capable of understanding/appreciating all the ways in which it's better.

Re: OpenAI Progress

#92
I just don't care about AGI.

I care a lot about AI coding.

OpenAI in particular seems to really think AGI matters. I don't think AGI is even possible because we can't define intelligence in the first place, but what do I know?

Re: OpenAI Progress

#93

I thought the response to "what would you say if you could talk to a future AI" would be "how many r in strawberry".

Can we stop with that outdated meme? What model can't answer that effectively?

GPT-5 can’t.

https://bsky.app/profile/kjhealy.co/post/3lvtxbtexg226

Re: OpenAI Progress

#94
post #5

What's really interesting is that if you look at "Tell a story in 50 words about a toaster that becomes sentient" (10/14), the text-davinci-001 is much, much better than both GPT-4 and GPT-5.

[deleted]

Re: OpenAI Progress

#95

Earlier quoted context omitted.

> I could essentially replace it with Google for basic to slightly complex fact checking. I know you probably meant "augment fact checking" here, but using LLMs for answering factual questions is the single worst use-case for LLMs.

I disagree. Some things are hard to Google, because you can't frame the question right. For example you know context and a poor explanation of what you are after. Googling will take you nowhere, LLMs will give you the right answer 95% of the time. Once you get an answer, it is easy enough to verify it.

If you’re looking for a possibly correct answer to an obscure question, that’s more like fact finding. Verifying it afterward is the “fact checking” step of that process.

Re: OpenAI Progress

#96

Earlier quoted context omitted.

> I could essentially replace it with Google for basic to slightly complex fact checking. I know you probably meant "augment fact checking" here, but using LLMs for answering factual questions is the single worst use-case for LLMs.

I disagree. Some things are hard to Google, because you can't frame the question right. For example you know context and a poor explanation of what you are after. Googling will take you nowhere, LLMs will give you the right answer 95% of the time. Once you get an answer, it is easy enough to verify it.

It's not the LLM alone though, it's “LLM with web search”, and as such 4o isn't really a leap at all there (IIRC perplexity was using an early Llama version and was already very good, long before OpenAI adding web search to ChatGPT).

Re: OpenAI Progress

#97

My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…

I have a theory about why it's so easy to underestimate long-term progress and overestimate short-term progress.

Before a technology hits a threshold of "becoming useful", it may have a long history of progress behind it. But that progress is only visible and felt to researchers. In practical terms, there is no progress being made as long as the thing is going from not-useful to still not-useful.

So then it goes from not-useful to useful-but-bad and it's instantaneous progress. Then as more applications cross the threshold, and as they go from useful-but-bad to useful-but-OK, progress all feels very fast. Even if it's the same speed as before.

So we overestimate short term progress because we overestimate how fast things are moving when they cross these thresholds. But then as fewer applications cross the threshold, and as things go from OK-to-decent instead of bad-to-OK, that progress feels a bit slowed. And again, it might not be any different in reality, but that's how it feels. So then we underestimate long-term progress because we've extrapolated a slowdown that might not really exist.

I think it's also why we see a divide where there's lots of people here who are way overhyped on this stuff, and also lots of people here who think it's all totally useless.

Re: OpenAI Progress

#98

My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…

the actual major leap was o1, going from 3.5 to 4 is just scaling, o1 is a different paradigm that skyrocketed its performance on math/physics problems (or reasoning more generally), it also made the model much more precise (essential for coding).

Re: OpenAI Progress

#99

Earlier quoted context omitted.

So it isn’t a “fact” that the built in Python function that tests whether a string ends with a substring is “endswith”? See https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect If you know that a source isn’t to be believed in an area you know about, why would you trust that source in an area you don’t know about? Another funny anecdote, ChatGPT just got the Gell-Man effect wrong. https://chatgpt.com/share/68a0b7af…

It got it right with thinking which was the challenge I posed. https://chatgpt.com/share/68a0b897-f8dc-800b-8799-9be2a8ad54...

The point you're missing is it's not always right. Cherry-picking examples doesn't really bolster your point.

Obviously it works for you (or at least you think it does), but I can confidently say it's fucking god-awful for me.

Post reply on HN