Live data from Hacker News

Generative AI's failure to induce robust models of the world

garymarcus.substack.com

51–60 of 89 posts

Re: Generative AI's failure to induce robust models of the world

#51
post #50

I usually disagree with Garry Marcus but his basic point seems fair enough if not surprising - Large Language Models model language about the world, not the world itself. For a human like understanding of the world you need some understanding of concepts like space, time, emotion, other creatures thoughts and so on, all things we pick up as kids. I don't see much reason why future AI couldn't do that rather than just…

The underlying assumption is that language and symbols are enough to represent phenomena. Maybe we are falling for this one in our own heads as well.

Understanding may not be a static symbolic representation. Contexts of the world infinite and continuously redefined. We believed we could represent all contexts tied to information, but that's a tough call.

Yes, we can approximate. No, we can't completely say we can represent every essential context at all times.

Some things might not be representable at all by their very chaotic nature.

Re: Generative AI's failure to induce robust models of the world

#52

Earlier quoted context omitted.

I think negative feedback loops of AIs trained on AI generated data might lead to a position where AI quality peaks and slides backwards.

I would not bet against synthetic data. AlphaZero is trained only on synthetic data and it's better than any human, and keeps getting better with more training compute. There is no negative feedback loop in the narrow cases we have tried previously. There may be trade-offs but on net we are going forward.

There's a pretty big difference between AlphaZero and a "generative AI" program: AlphaZero has access to an oracle that can tell it whether it's making valid moves and winning games.

By comparison, getting accurate feedback on whether facts are correct in a piece of text (for example) is much more difficult and expensive. At least, presumably that's why AI companies publish staged demo videos where the AI still makes factual errors half the time.

Re: Generative AI's failure to induce robust models of the world

#53
post #51
post #50

I usually disagree with Garry Marcus but his basic point seems fair enough if not surprising - Large Language Models model language about the world, not the world itself. For a human like understanding of the world you need some understanding of concepts like space, time, emotion, other creatures thoughts and so on, all things we pick up as kids. I don't see much reason why future AI couldn't do that rather than just…

The underlying assumption is that language and symbols are enough to represent phenomena. Maybe we are falling for this one in our own heads as well. Understanding may not be a static symbolic representation. Contexts of the world infinite and continuously redefined. We believed we could represent all contexts tied to information, but that's a tough call. Yes, we can approximate. No, we can't completely say we can re…

I did think that human mental modeling of the world is also quite rough and often inaccurate. I don't see why AI can't become human like in it's abilities but accurately modeling all the relativistic quarks in an atom is a bit beyond anything just now.

Re: Generative AI's failure to induce robust models of the world

#54

The whole thing is silly. Look, we know that LLMs are just really good word predictors. Any argument that they are thinking is essentially predicated on marketing materials that embrace anthropomorphic metaphors to an extreme degree. Is it possible that reason could emerge as the byproduct of being really good at predicting words? Maybe, but this depends on the antecedent claim that much if not all of reason is stric…

What use of the word "reasoning" are you trying to claim that current language models knowably fail to qualify for, except that it wasn't done by a human?

Well - all of them.

The mechanism by which they work prohibits reasoning.

This is easy to see if you look at a transformer architecture and think through what each step is doing.

The amazing thing is that they produce coherent speech, but they literally can't reason.

Re: Generative AI's failure to induce robust models of the world

#55
post #45

Earlier quoted context omitted.

Yeah, that's fair. It's probably more accurate to call them sequence predictors or general data predictors than to limit it to words (unless we mean words in the broad, mathematical sense) they are free monoid emulators

And what are humans?

Humans are humans - to deny that we are thinking, reasoning, living beings is a strange thing to do.

You can taste a beer, laugh so much it hurts, come to know how something works.

Re: Generative AI's failure to induce robust models of the world

#56
post #40

Earlier quoted context omitted.

The lack of information in ant trails (beyond "it exists here") leads to death spirals https://en.m.wikipedia.org/wiki/Ant_mill

The very first sentence of the article you linked says this happens because they lose the pheromone track.

It does, but the original paper is not as certain. (https://digitallibrary.amnh.org/server/api/core/bitstreams/8...) It suggests that the following is done "both chemically and tactually" in case of a circular path, although with just a minimal test.

Re: Generative AI's failure to induce robust models of the world

#57

I definitely would be okay if we hit an AI winter; our culture and world cannot adapt fast enough for the change we are experiencing. In the meantime, the current level of AI is just good enough to make us more productive, but not so good as to make us irrelevant.

I think negative feedback loops of AIs trained on AI generated data might lead to a position where AI quality peaks and slides backwards.

We are just at the beginning of integrating external tools in the process and developing complex cognitive structures. LLM is just one part of it. Till now it was cheaper and easier to improve that part especially if other work would be rendered obsolete by LLM improvements.

Re: Generative AI's failure to induce robust models of the world

#58

Earlier quoted context omitted.

I would not bet against synthetic data. AlphaZero is trained only on synthetic data and it's better than any human, and keeps getting better with more training compute. There is no negative feedback loop in the narrow cases we have tried previously. There may be trade-offs but on net we are going forward.

There's a pretty big difference between AlphaZero and a "generative AI" program: AlphaZero has access to an oracle that can tell it whether it's making valid moves and winning games. By comparison, getting accurate feedback on whether facts are correct in a piece of text (for example) is much more difficult and expensive. At least, presumably that's why AI companies publish staged demo videos where the AI still makes…

Automatic verification (oracle) is being used today to create synthetic data for LLMs. I don't see it as a big difference versus AlphaZero. While there's no way to ensure that a single synthetic reasoning trace is correct, as long as it leads to the correct answer according to the verifier, the law of large numbers should take care of that.

The problem is that it's difficult to create verifiers for many things we care about like architectural taste. So I expect to see superhuman capabilities on the things we can make verifiers for, but for other things it's harder to predict. We may see transfer learning or we may see collapse. My money would be more on transfer learning.

Re: Generative AI's failure to induce robust models of the world

#59
post #18

Earlier quoted context omitted.

I don't understand the reasoning behind drawing a conclusion that if something fails a task that requires reasoning implies that thing cannot reason. To use chess as an example. Humans sometimes play illegal moves. That does not mean Humans cannot reason. It is an instance of failing to show proof of reasoning. Not a proof of the inability to reason.

Anthropomorphic fallacy. Human fails at task due to not knowing the rules in perfect detail. AI fails at task even though it knows the rules and could easily reproduce them for chess and dozens of chess variants. "Look! The fallibility of humans rubbed off onto the AI, proving that they are more human and AGI than we give them credit to!"

I'm not sure how you consider this to be an anthropomorphic fallacy, the comparison to the situation with a human exists only because people are prepared to stipulate that humans can reason. That does not assume something about AI behaviour to be like a human's. It is showing the same test applied to a human.

Your statement that AI knows the rules would be considered anthropomorphising by many, I take it more to mean it 'knows' in the same sense that an election 'wants' to be at a lower energy level.

That said, humans who have written entire books on chess have been known to play illegal moves. That should count as proof by counterexample that your reasoning as to why humans fail at tasks is false.

Re: Generative AI's failure to induce robust models of the world

#60
post #32

That LLMs are a black box and that LLMs lack an underlying model are both true, but orthogonal. It's possible to have a black box system which has an underlying model. That's true of many statistical prediction methods. Early attempts at machine learning were a white box with no underlying model. This is true of most curve-fitting. The AI version was where you're trying to divide a high-dimensional space with a cutti…

“LLMs lack an underlying model” is very obviously incorrect. LLMs have an underlying model of semantics as tokens embedded into a high-dimensional vector space. The question is not whether or not they have any model at all, the question is whether the model they indisputably have (which is a model of language in terms of linear algebra) maps onto a model of the external universe (a “world model”) that emerges during…

I agree that it's probably unfalsifiable in the sense of proving it definitively based on something like static analysis of the model itself.

But that doesn't mean that we can't, in theory, give the LLM a battery of tests that it should perform well (though not perfectly) on if it has a world model, and poorly (though not fail totally) on if it doesn't.

It's inherently a probabilistic system, so testing it in a probabilistic manner seems perfectly apt. Again: no, this will not produce a definitive result, due to that probabilistic nature—but it can produce an indicative one, and running the same test on multiple related LLMs, or similar tests on the same LLM, should help to smooth out noise in the results.

(...of course, this only works if the tests are designed well, and I don't have enough specific understanding of LLMs to know how one would go about doing that in a rigorous manner!)

Post reply on HN