Live data from Hacker News

OpenAI Progress

progress.openai.com

311–320 of 372 posts

Re: OpenAI Progress

#311
post #280
post #134

Earlier quoted context omitted.

They outperform asking humans, unless you are asking an expert. On average

When I have a question, I don't usually "ask" that question and expect an answer. I figure out the answer. I certainly don't ask the question to a random human.

you ask yourself .. for most people, that means closer to average reply, from yourself, when you try to figure it out.

There is a working paper from McKinnon Consulting in Canada that states directly that their definition of "General AI" is when the machine can match or exceed fifty percent of humans who are likely to be employed for a certain kind of job. It implies that low-education humans are the test for doing many routine jobs, and if the machine can beat 50% (or more) of them with some consistency, that is it.

Re: OpenAI Progress

#312
post #243

Earlier quoted context omitted.

At the time getting complete sentences was extremely difficult! N-gram models were essentially the best we had

No, it was not difficult at all. I really wonder why they have such a bad example here for GPT1. See for example this popular blog post: https://karpathy.github.io/2015/05/21/rnn-effectiveness/ That was in 2015, with RNN LMs, which are all much much weaker in that blog post compared GPT1. And already looking at those examples in 2015, you could maybe see the future potential. But no-one was thinking that scaling up w…

Thanks for posting some of the history... "You might find even earlier examples" is pretty tongue-in-cheek though. [1], expanded in 2003 into [2], has 12466 citations, 299 by 2011 (according to Google Scholar which seems to conflate the two versions). The abstract [2] mentions that their "large models (with millions of parameters)" "significantly improves on state-of-the-art n-gram models, and... allows to take advantage of longer contexts." Progress between 2000 and 2017 (transformers) was slow and models barely got bigger.

And what people forget about Mikolov's word2vec (2013) was that it actually took a huge step backwards from the NNs like [1] that inspired it, removing all the hidden layers in order to be able to train fast on lots of data.

[1] Yoshua Bengio, Réjean Ducharme, Pascal Vincent, 2000, NIPS, A Neural Probabilistic Language Model

[2] Yoshua Bengio, Réjean Ducharme, Pascal Vincent, Christian Jauvin, 2003, JMLR, A Neural Probabilistic Language Model, https://www.jmlr.org/papers/volume3/bengio03a/bengio03a.pdf

Re: OpenAI Progress

#313

This feels like Flowers for Algernon

Yeah. And like watching a child grow up.

no disagree and object -- it is misinformation to describe LLM progress blindly assigning human development. Tons of errant conclusions and implications attached as baggage, and just wrong.

Re: OpenAI Progress

#314

Earlier quoted context omitted.

I'd love to know more about how OpenAI (or Alec Radford et al.) even decided GPT-1 was worth investing more into. At a glance the output is barely distinguishable from Markov chains. If in 2018 you told me that scaling the algorithm up 100-1000x would lead to computers talking to people/coding/reasoning/beating the IMO I'd tell you to take your meds.

There's a performance plateau with training time and number of parameters and then once you get over "the hump" error rate starts going down again almost linearly. GPT existed before OpenAI but it was theorized that the plateau was a dead end. The sell to VCs in the early gpt3 era was "with enough compute, enough time, and enough parameters... it'll probably just start thinking and then we have AGI". Sometime around…

source? i work in this field and have never heard of the initial plateau you are referring

Re: OpenAI Progress

#315

Earlier quoted context omitted.

All the replies are spectacularly wrong, and biased by hindsight. GPT-1 to GPT-2 is where we went from "yes, I've seen Markov chains before, what about them?" to "holy shit this is actually kind of understanding what I'm saying!" Before GPT-2, we had plain old machine learning. After GPT-2, we had "I never thought I would see this in my lifetime or the next two".

GPT-2 was the first wake-up call - one that a lot of people slept through. Even within ML circles, there was a lot of skepticism or dismissive attitudes about GPT-2 - despite it being quite good at NLP/NLU. I applaud those who had the foresight to call it out as a breakthrough back in 2019.

i think it was already pretty clear among practitioners by 2018 at the latest

Re: OpenAI Progress

#317

I’m baffled by claims that AI has “hit a wall.” By every quantitative measure, today’s models are making dramatic leaps compared to those from just a year ago. It’s easy to forget that reasoning models didn’t even exist a year back! IMO Gold, Vibe coding with potential implications across sciences and engineering? Those are completely new and transformative capabilities gained in the last 1 year alone. Critics argue…

I don't think it is that surprising.

It will become harder and harder for the average person to gain from newer models.

My 75 year old father loves using Sonnet. He is not asking anything though that he would be able to tell Opus is "better". The answers he gets from the current model are good enough. He is not exactly using it to probe the depths of statistical mechanics.

My father is never going to vibe code anything no matter how good the models get.

I don't think AGI would even give much different answers to what he asks.

You have to ask the model something that allows the latest model to display its improvements. I think we can see, that is just not something on the mind of the average user.

Re: OpenAI Progress

#318

Earlier quoted context omitted.

I have a theory about why it's so easy to underestimate long-term progress and overestimate short-term progress. Before a technology hits a threshold of "becoming useful", it may have a long history of progress behind it. But that progress is only visible and felt to researchers. In practical terms, there is no progress being made as long as the thing is going from not-useful to still not-useful. So then it goes from…

Your threshold theory is basically Amara's Law with better psychological scaffolding. Roy Amara nailed the what ("we tend to overestimate the effect of a technology in the short run and underestimate the effect in the long run") [1] but you're articulating the why better than most academic treatments. The invisible-to-researchers phase followed by the sudden usefulness cascade is exactly how these transitions feel fr…

One thing I think is weird in the debate is it seems people are equating LLMs with CPU’s, this whole category of devices that process and calculate and can have infinite architecture and innovation. But what if LLMs are more like a specific implementation like DSP’s, sure lots of interesting ways to make things sound better, but it’s never going to fundamentally revolutionize computing as a whole.

Re: OpenAI Progress

#319

Earlier quoted context omitted.

Why? It sounds like you're using "I believe it's rapidly getting smarter" as evidence for "so it's getting smarter in ways we don't understand", but I'd expect the causality to go the other way around.

Simply because of what we know about our ability to judge capabilities and systems. It's much harder to judge solutions to hard problems. You can demonstrate that you can add 2+2, and anyone* can be the judge of that ability, but if you try to convince anyone of a mathematical proof you came up with, that would be a much harder thing to do, regardless of your capability to write that prove and how hard it was to writ…

> The more complicated and/or complex things become, the less likely it is that a human can act as a reliable judge. At some point no human can.

Give me an example, please. I can't come up with something that started simple and became too complex for humans to "judge". I am quite curious.

Re: OpenAI Progress

#320

Everyone knows they edited these right to show the progression they wanted right?

It seems more likely that they cherry-picked favorable examples showing a clear progression, rather than just falsifying the content.

Even in these comments, there's a fair bit of disagreement about whether they show do show monotonic improvement.

Post reply on HN