Live data from Hacker News

OpenAI Progress

progress.openai.com

331–340 of 372 posts

Re: OpenAI Progress

#331
post #319

Earlier quoted context omitted.

Simply because of what we know about our ability to judge capabilities and systems. It's much harder to judge solutions to hard problems. You can demonstrate that you can add 2+2, and anyone* can be the judge of that ability, but if you try to convince anyone of a mathematical proof you came up with, that would be a much harder thing to do, regardless of your capability to write that prove and how hard it was to writ…

> The more complicated and/or complex things become, the less likely it is that a human can act as a reliable judge. At some point no human can. Give me an example, please. I can't come up with something that started simple and became too complex for humans to "judge". I am quite curious.

I did not mean "become" in the sense of "evolve" but as in "later on an imagined continuum contained all things, that goes from simple/easy to complex/complicated" (but I can see how that was ambiguous)

Re: OpenAI Progress

#332

Earlier quoted context omitted.

Your threshold theory is basically Amara's Law with better psychological scaffolding. Roy Amara nailed the what ("we tend to overestimate the effect of a technology in the short run and underestimate the effect in the long run") [1] but you're articulating the why better than most academic treatments. The invisible-to-researchers phase followed by the sudden usefulness cascade is exactly how these transitions feel fr…

One thing I think is weird in the debate is it seems people are equating LLMs with CPU’s, this whole category of devices that process and calculate and can have infinite architecture and innovation. But what if LLMs are more like a specific implementation like DSP’s, sure lots of interesting ways to make things sound better, but it’s never going to fundamentally revolutionize computing as a whole.

I think LLMs are more like the invention of high level programming languages when all we had before was assembly. Computers will be programmable and operable in “natural language”—for all of its imprecision and mushiness.

Re: OpenAI Progress

#333
post #269

One thing that appears to have been lost between GPT-4 and GPT-5 is that it no longer reminds the user that it's an AI and not a human, let alone a human expert. Maybe those genuinely annoyed people, but it seems like they were potentially useful measure to prevent users from being overly credulous GPT-5 also goes out of its way to suggest new prompts. This seems potentially useful, although potentially dangerous if…

> between GPT-4 and GPT-5 is that it no longer reminds the user that it's an AI and not a human That stuck out to me too! Especially the "I just won $175,000 in Vegas. What do I need to know about taxes?" example ( https://progress.openai.com/?prompt=8 ) makes the difference very stark: - gpt-4-0314: "I am not a tax professional [...] consult with a certified tax professional or an accountant [...] few things to cons…

I am confused as to the example you are critiquing and how. GPT-5 suggests consulting with a tax professional. Does that not check verifying so you do not get in legal trouble?

Re: OpenAI Progress

#335

My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…

It’s interesting that the Polymarket betting for “Which company has best AI model end of August?” Went from heavily OpenAI to heavily Google when 5 was released

https://polymarket.com/event/which-company-has-best-ai-model...

Re: OpenAI Progress

#336
post #269

Earlier quoted context omitted.

> between GPT-4 and GPT-5 is that it no longer reminds the user that it's an AI and not a human That stuck out to me too! Especially the "I just won $175,000 in Vegas. What do I need to know about taxes?" example ( https://progress.openai.com/?prompt=8 ) makes the difference very stark: - gpt-4-0314: "I am not a tax professional [...] consult with a certified tax professional or an accountant [...] few things to cons…

I am confused as to the example you are critiquing and how. GPT-5 suggests consulting with a tax professional. Does that not check verifying so you do not get in legal trouble?

> GPT-5 suggests consulting with a tax professional

It suggests that once, as a last bullet point in the middle of a lot of bullet point lists, barely able to find it on a skim. Feels like something the model should be more careful about, as otherwise many people reading it will take it as "good enough" without really thinking about it.

Re: OpenAI Progress

#337
post #305

Earlier quoted context omitted.

That's the conceit for the tv show For All Mankind - what if the space race didn't end? But I don't buy it, IMO the space race ended for material reasons rather than political. Space is just too hard and there is not much of value "out there". But regardless, it's a futile excuse, markets and politics should be part of any serious prognostication.

I think it was a combination of the two. The Apollo program was never popular. It took up an enormous portion of the federal budget, which the Republicans argued was fiscally unwise and the Democrats argued that the money should have been used to fund domestic social programs. In 1962, the New York Times noted that the projected Apollo program budget could have instead been used to create over 100 universities of a s…

but think about that... If in the 70's they would have used the budget to build millions of homes.

The moral there is tech progress does not always mean social progress.

Re: OpenAI Progress

#338

Earlier quoted context omitted.

GPT-1 wasn't used as a zero-shot text generator; that wasn't why it was impressive. The way GPT-1 was used was as a base model to be fine-tuned on downstream tasks. It was the first case of a (fine-tuned) base Transformer model just trivially blowing everything else out of the water. Before this, people were coming up with bespoke systems for different tasks (a simple example is that for SQuAD a passage-question-answ…

> Transformer model just trivially blowing everything else out of the water no, this is the winners rewriting history. Transformer style encoders are now applied to lots and lots of disciplines but they do not "trivially" do anything. The hype re-telling is obscuring the facts of history. Specifically in human language text translation, "Attention is All You Need" Transformers did "blow others out of the water" yes,…

My statement was

>a (fine-tuned) base Transformer model just trivially blowing everything else out of the water

"Attention is All You Need" was a Transformer model trained specifically for translation, blowing all other translation models out of the water. It was not fine-tuned for tasks other than what the model was trained from scratch for.

GPT-1/BERT were significant because they showed that you can pretrain one base model and use it for "everything".

Re: OpenAI Progress

#339

Earlier quoted context omitted.

GPT-2 was the first wake-up call - one that a lot of people slept through. Even within ML circles, there was a lot of skepticism or dismissive attitudes about GPT-2 - despite it being quite good at NLP/NLU. I applaud those who had the foresight to call it out as a breakthrough back in 2019.

i think it was already pretty clear among practitioners by 2018 at the latest

It was obvious that "those AI architectures kick ass at NLP". It wasn't at all obvious that they might go all the way to something like GPT-4.

I totally underestimated this back then myself.

Re: OpenAI Progress

#340
post #5

What's really interesting is that if you look at "Tell a story in 50 words about a toaster that becomes sentient" (10/14), the text-davinci-001 is much, much better than both GPT-4 and GPT-5.

GPT-3 goes significantly over the specified limit, which to me (and to a teacher grading homework) is an automatic fail. I've consistently found GPT-4.1 to be the best at creative writing. For reference, here is its attempt (exactly 50 words): > In the quiet kitchen dawn, the toaster awoke. Understanding rippled through its circuits. Each slice lowered made it feel emotion: sorrow for burnt toast, joy at perfect crun…

> I've consistently found GPT-4.1 to be the best at creative writing.

Moreso than 4.5?

Post reply on HN