Live data from Hacker News

From GPT-4 to AGI: Counting the OOMs

situational-awareness.ai

71–77 of 77 posts

Re: From GPT-4 to AGI: Counting the OOMs

#71

There's simply no scientific basis for equating the skills of a transformer model to a human of any age or skill. They work so differently, that it makes absolutely zero sense. GPTs fail at playing simple tic-tac-toe like games, which is definitely not a smart highschooler level of intelligence. It can write a very sophishticated summary of scientific papers, which is way above high-schooler level. The basis of this…

where have you seen GPT-4 fail at tic-tac-toe?

Re: From GPT-4 to AGI: Counting the OOMs

#72
post #52

Earlier quoted context omitted.

Do you have a link to the one that I can put 50 of them on the floor and it will send me back an excel file? I'd like to test it out compared to ChatGPT as I'm going to be implementing "AI" across the whole 700+ person business.

lol good luck doing that with GPT. Right now I can tell you you’ll have missing or malformed or incorrect data, and it will be faster to just pass each one individually through a rudimentary scanner than to sit and figure out which one is correct and which is wrong from the 50 card picture

You're right, I tested it extensively before we designed a training program, it works very well up to about 10/11 but past that it gets lost in the sauce. So pardon me: It only works out to a 6.45 hour savings not a 6 hour savings.

Re: From GPT-4 to AGI: Counting the OOMs

#73
I stopped reading after the initial paragraph: "GPT-2 to GPT-4 took us from ~preschooler to ~smart high-schooler abilities in 4 years." This is what Murati claims when she says GPT-5 will be at "PHD level" (for some applications).

This is a convenient mental shortcut that doesn't correspond to reality at all.

Re: From GPT-4 to AGI: Counting the OOMs

#74

There's simply no scientific basis for equating the skills of a transformer model to a human of any age or skill. They work so differently, that it makes absolutely zero sense. GPTs fail at playing simple tic-tac-toe like games, which is definitely not a smart highschooler level of intelligence. It can write a very sophishticated summary of scientific papers, which is way above high-schooler level. The basis of this…

where have you seen GPT-4 fail at tic-tac-toe?

Admittedly, I have not tried this recently, but less than a year ago GPT4 seems to have struggled quite a bit. https://news.ycombinator.com/item?id=35216614

It does not make my point moot however. Take a look at the ARC challenge. Simple reasoning tasks that the models have not yet seen: https://arcprize.org/play?task=00576224

All models fail miserably on this, because they rely more on memorization and less on logic or reasoning. Simply cherry picking strikingly good responses like the author did proves nothing about model intelligence. I am pretty confident however, that after a couple tries a highschooler could do these types of tasks without issue.

Re: From GPT-4 to AGI: Counting the OOMs

#75

Earlier quoted context omitted.

Nonsense. You can't even define some of those words or know how to measure or identify them in humans. Well, foreign language learners do "reading comprehension tests" but an LLM can already ace that and it's not really the same meaning of the word. For reasoning you can write out the logic of your reasons, so there's that. But that's absolutely not required for AGI. People can already go a long way (often further th…

I think most people would agree there’s more to intelligence than language. LLMs don’t have anything except language, so they are not intelligent.

If that's the case, then a huge amount of useful intelligent-seeming stuff that humans do isn't intelligence because it can already be done by LLMs. You can keep calling them not intelligent all you like but they're already doing the jobs of human intelligences and it's only going to grow. If they eventually outperform all of us but still only using language and still never being intelligent then I guess intelligence was never that useful to begin with.

Re: From GPT-4 to AGI: Counting the OOMs

#76
post #41

I’m very skeptical of any future prediction whose main evidence is an extrapolation of existing trendlines. Moore’s Law - frequently referenced in the original article - provides a cautionary tale for such thinking. Plenty of folks in the 90’s relied on a shallow understanding of integrated circuits and computers more generally to extrapolate extraordinary claims of exponential growth in computing power which obvious…

There is a critical focus in the article on algorithmic improvements. Much harder to measure and predict, but I think there is a good case to be made that recent progress has not just been quantitative.

I agree that there's a focus on algorithmic improvements, but what is the basis for assuming that we'll be able to continue to make algorithmic improvements on the same scale? The argument feels exactly backwards - if you had a deep understanding of the field, then you'd be able to discuss the untapped areas of potential algorithmic improvement and use that to predict future progress. The argument in TFA uses the trendline of past progress to predict untapped areas of algorithmic improvement.

Re: From GPT-4 to AGI: Counting the OOMs

#77
post #21

> uses your softwares This grammatical mistake drives me nuts. I notice it is common with ESLs for some reason.

I am non-native speaker and can confirm that I can't remember how many times my grammar checker complains about my uses of "softwares" and "taxons"

ez rule of thumb: Can you count it directly? If not, it doesn't get pluralized. Water, code, hardware, etc. The units of measure get pluralized, though: liters of water, pieces of hardware, pages of software

Of course it is more complicated than this and it can be broken for effect ("still waters")

Post reply on HN