Live data from Hacker News

Sparks of Artificial General Intelligence: Early Experiments with GPT-4

arxiv.org

51–60 of 244 posts

Re: Sparks of Artificial General Intelligence: Early Experiments with GPT-4

#51
This is a pretty fluffy paper, especially for an institution like Microsoft Research. It says it's an "early AGI" in the abstract, but elsewhere says it's merely a "step towards AGI". The basis for this is asking ChatGPT a bunch of stuff, but they don't really present an overarching framework for what questions to ask or why.

The paper makes outlandish claims like "GPT-4 has common sense grounding" on the basis of its answers to these questions, but the questions don't show that the model has common sense or grounding. One of their constructed questions involves prompting the model with the equator's exact length—"precisely 24,901 miles"—and then being astonished that the model predicts that you're on the equator ("Equator" being the first result on Wikipedia for the search term "24,901"). It's also the case that while GPT-4 can say a bear at the north pole is "white", it has no way of knowing what "white", or "bear", or "north" actually represent.

Are there folks out there doing rigorous research on these topics, who have a framework for developing tests of actual understanding?

Re: Sparks of Artificial General Intelligence: Early Experiments with GPT-4

#52
post #14

ChatGPT and its relatives are very very impressive on first impressions, but I've been using ChatGPT-3 and now 4 heavily every day since they became available to individuals and once you start using them this much it becomes very clear how NOT intelligent they are. It really just seems like extremely impressive statistical inference after this much use and finding so many failure modes. But it is still impressive how…

It depends on 1) the domains 2) your comparison group.

On 2), many software engineers and computer scientists compare these language models' logic and creative problem solving abilities with themselves and their peer group. But they are usually 1-2+ SD above average humans at these things.

(Note: Someone gave GPT-4 an IQ test and the result was 96, slightly below the average of reference human group at 100. The SD of an IQ test is 15 or 16.)

For language-focused domains, there is evidence that GPT-4 is already better than most humans, eg. 99th percentile at GRE Verbal, beat humans at a fairly novel puzzle like Twofer Goofer, which is not in its training set.

Ref: GPT-4 Beats Humans at Hard Rhyme-based Riddles https://twofergoofer.com/blog/gpt-4

Yes, GPT-4 is not an AGI yet, but the research paper (OP) has a point.

Re: Sparks of Artificial General Intelligence: Early Experiments with GPT-4

#53
post #51

This is a pretty fluffy paper, especially for an institution like Microsoft Research. It says it's an "early AGI" in the abstract, but elsewhere says it's merely a "step towards AGI". The basis for this is asking ChatGPT a bunch of stuff, but they don't really present an overarching framework for what questions to ask or why. The paper makes outlandish claims like "GPT-4 has common sense grounding" on the basis of it…

>it has no way of knowing what "white", or "bear", or "north" actually represent.

What does it mean to know what "white", "bear" or "north" actually represent?

Re: Sparks of Artificial General Intelligence: Early Experiments with GPT-4

#54

Earlier quoted context omitted.

How do you believe they're orthogonal?

Biological reproduction for one. Copy/paste for another. Biological reproduction is only tenuously related to understanding and copy/paste isn’t even related at all. We can copy around weights and biases all day without understanding them.

Makes sense, you've changed my mind a bit but only the part that consciously understands.

Re: Sparks of Artificial General Intelligence: Early Experiments with GPT-4

#55
post #22
post #14

ChatGPT and its relatives are very very impressive on first impressions, but I've been using ChatGPT-3 and now 4 heavily every day since they became available to individuals and once you start using them this much it becomes very clear how NOT intelligent they are. It really just seems like extremely impressive statistical inference after this much use and finding so many failure modes. But it is still impressive how…

Our of curiosity, what is GPT-4 getting wrong so often? It’s prettily wild to my own , admittedly easily impressed, mind.

GPT is really good at repeating what the average intelligent response to something might look like, but it doesn't seem to be actually reasoning about any of its responses. Give it a complex logical problem that it needs to deduce from inputs, such as which foods contain gluten, based on their ingredient lists, and it will reliably fail. As a person with celiac, this is a task I complete multiple times a day with no effort. Just today I was trying to build a prompt that would summarize daily news updates leaving out anything about Russia, but it still included Russia more often than not despite being very clear in the prompt that anything about Russia should not being included in the response under any circumstances.

Re: Sparks of Artificial General Intelligence: Early Experiments with GPT-4

#56

I remember reading this somewhere - "There is a considerable overlap between the intelligence of the smartest bears and the dumbest tourists.". Though I do not think GPT-4 is even close to AGI it can definitely claim to be better at faking it than many intelligent beings can.

so we are at the snapshot in time where people think 'AI is smarter than many people but not even close to being as smart as me'

Re: Sparks of Artificial General Intelligence: Early Experiments with GPT-4

#57
post #13

> Given the breadth and depth of GPT-4's capabilities, we believe that it could reasonably be viewed as an early (yet still incomplete) version of an artificial general intelligence (AGI) system. But it's just statistics, a fancy text predictor, a Markov-chain. Surely these scientists that work in the field of AI and are intimately familiar with how this stuff works aren't so stupid as to think emergent behavior pote…

GPT-4 is often overhyped and underhyped because few really understand it. It's not a Markov Chain or a fancy text predictor. It's a ~200 layer neural network that models a vast hierarchy of concepts through language. It has emergent properties that we don't yet understand.

Where are you getting the 200 number from?

Re: Sparks of Artificial General Intelligence: Early Experiments with GPT-4

#59

Earlier quoted context omitted.

GPT-4 is often overhyped and underhyped because few really understand it. It's not a Markov Chain or a fancy text predictor. It's a ~200 layer neural network that models a vast hierarchy of concepts through language. It has emergent properties that we don't yet understand.

Where are you getting the 200 number from?

I must have hallucinated that. GPT-3 has 96 layers but they haven't disclosed the number of layers in GPT-4.
Post reply on HN