Live data from Hacker News

OpenAI Progress

progress.openai.com

321–330 of 372 posts

Re: OpenAI Progress

#321

My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…

I have a theory about why it's so easy to underestimate long-term progress and overestimate short-term progress. Before a technology hits a threshold of "becoming useful", it may have a long history of progress behind it. But that progress is only visible and felt to researchers. In practical terms, there is no progress being made as long as the thing is going from not-useful to still not-useful. So then it goes from…

GPT3 is when the mass started to get exposed to this tech, it felt like a revolution.

Got 3.5 felt like things were improving super super fast and created that feeling the near feature will be unbelievable.

Got to 4/o series, it felt things had improved but users weren't as thrilled as with the leap to 3.5

You can call that bias, but clearly version 5 improvements displays an even greater slow down, that's 2 long years since gp4.

For context:

- gpt 3 got out in 2020

- gpt 3.5 in 2022

- gpt 4 in 2023

- gpt 4o and clique, 2024

After 3.5 things slowed down, in term of impact at least. Larger context window, multi-modality, mixture of experts, and more efficienc: all great, significant features, but all pale compared to the impact made by RLHF already 4 years ago.

Re: OpenAI Progress

#322

Earlier quoted context omitted.

Every time I use ChatGPT I become incredibly frustrated with how fucking awful it is. I've used it more than enough, time and time again (just try the new model, bro!), to know that I fucking hate it. If it works for you, cool. I think it's dogshit.

They just spent like six comments imploring you to understand that they were making a specific point: generally reliable on non-niche topics using thinking mode. And that nuance bounced off of you every single time as you keep repeating it's not perfect, dismiss those qualifications as cherry picking and repeat personal anecdotes. I'm sorry but this is a lazy and unresponsive string of comments that's degrading the d…

The neat thing about HN is we can all talk about stupid shit and disagree about what matters. People keep upvoting me, so I guess my thoughts aren't unpopular and people think it's adding to the discussion.

I agree this is a stupid comment thread, we just disagree about why.

Re: OpenAI Progress

#323
post #57
post #5

What's really interesting is that if you look at "Tell a story in 50 words about a toaster that becomes sentient" (10/14), the text-davinci-001 is much, much better than both GPT-4 and GPT-5.

GPT 4.5 (not shown here) is by far the best at writing.

Aren't they discontinuing 4.5 in favor of 4.1? I think they already have with the API.

Re: OpenAI Progress

#324

Geez! When it comes to answering questions, GPT-5 almost always starts with glazing about what a great question it is, where as GPT-4 directly addresses the answer without the fluff. In a blind test, I would probably pick GPT-4 as a superior model, so I am not surprised why people feel so let down with GPT-5.

GPT-4 starts many responses with "As an AI language model", "I'm an AI", "I am not a tax professional", "I am not a doctor". GPT-5 does away with that and assumes an authoritative tone.

That makes fundraising easier, by increasing the appearance of authority and therefore coming off as a "better" model. In the elo ratings I'm sure GPT-5 is doing better because of the clear push back against these LLMs as ways to "cheat" without understanding and to flood propaganda -- better not to mention it.

Re: OpenAI Progress

#325

I’m baffled by claims that AI has “hit a wall.” By every quantitative measure, today’s models are making dramatic leaps compared to those from just a year ago. It’s easy to forget that reasoning models didn’t even exist a year back! IMO Gold, Vibe coding with potential implications across sciences and engineering? Those are completely new and transformative capabilities gained in the last 1 year alone. Critics argue…

I don't think it is that surprising. It will become harder and harder for the average person to gain from newer models. My 75 year old father loves using Sonnet. He is not asking anything though that he would be able to tell Opus is "better". The answers he gets from the current model are good enough. He is not exactly using it to probe the depths of statistical mechanics. My father is never going to vibe code anythi…

Correct. People claim these models "saturate" yet what saturates faster is our ability to grasp what these models are capable of.

I, for one, cannot evaluate the strength of an IMO gold vs IMO bronze models.

Soon coding capabilities might also saturate. It might all become a matter of more compute (~ # iterations), instead of more precision (~ % getting it right the first time), as the models become lightning speed, and they gain access to a playground.

Re: OpenAI Progress

#326
post #298

My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…

Everyone talks about 4o so positively but I’ve never consistently relied on it in a production environment. I’ve found it to be inconsistent in json generation and often it’s writing and following of the system prompt was very poor. In fact it was a huge part of what got me looking closer at anthropics models. I’m really curious what people did with it because while it’s cool it didn’t compare well in my real world u…

For json generation (and most API things) you should be using “structured outputs”

Re: OpenAI Progress

#328

Earlier quoted context omitted.

With the exception of h-bomb/fusion and ENIAC/AI, I think all of those examples reflect a change in priority and investment more than anything. There was a trajectory of high investment / rapid progress, then market and social and political drivers changed and space travel / supersonic flight just became less important.

That's the conceit for the tv show For All Mankind - what if the space race didn't end? But I don't buy it, IMO the space race ended for material reasons rather than political. Space is just too hard and there is not much of value "out there". But regardless, it's a futile excuse, markets and politics should be part of any serious prognostication.

[deleted]

Re: OpenAI Progress

#329

Earlier quoted context omitted.

That's the conceit for the tv show For All Mankind - what if the space race didn't end? But I don't buy it, IMO the space race ended for material reasons rather than political. Space is just too hard and there is not much of value "out there". But regardless, it's a futile excuse, markets and politics should be part of any serious prognostication.

I think the space race ended because we got all the benefit available, which wasn’t really in space anyway, it was the ancillary technical developments like computers, navigation, simulation, incredible tolerances in machining, material science, etc. We’re seeing a resurgence in space because there is actually value in space itself, in a way that scales beyond just telecom satellites. Suddenly there are good reasons…

> I think the space race ended because we got all the benefit available

We did get all the things that you listed but you missed the main reason it was started: military superiority. All of the other benefits came into existence in service of this goal.

Re: OpenAI Progress

#330
Huh. I find myself preferring aspects of TEXT-DAVINCI-001 in just about every example prompt.

It’s so to-the-point. No hype. No overstepping. Sure, it lacks many details that later models add. But those added details are only sometimes helpful. Most of the time, they detract.

Makes me wonder what folks would say if you re-released TEXT-DAVINCI-001 as “GPT5-BREVITY”

I bet you’d get split opinions on these types of not so hard / creative prompts.

Post reply on HN