Live data from Hacker News

Generative AI's failure to induce robust models of the world

garymarcus.substack.com

71–80 of 89 posts

Re: Generative AI's failure to induce robust models of the world

#71

I definitely would be okay if we hit an AI winter; our culture and world cannot adapt fast enough for the change we are experiencing. In the meantime, the current level of AI is just good enough to make us more productive, but not so good as to make us irrelevant.

[dead]

Re: Generative AI's failure to induce robust models of the world

#72
post #38

Earlier quoted context omitted.

Thank goodness we have version control systems then.

"Version control systems", in case of AI, mean that their knowledge will stay frozen in time, and so their usefulness will diminish. You need fresh data to train AI systems on, and since contemporary data is contaminated with generative AI, it will inevitably lead to inbreeding and eventual model collapse.

[dead]

Re: Generative AI's failure to induce robust models of the world

#73
post #18

Earlier quoted context omitted.

I don't understand the reasoning behind drawing a conclusion that if something fails a task that requires reasoning implies that thing cannot reason. To use chess as an example. Humans sometimes play illegal moves. That does not mean Humans cannot reason. It is an instance of failing to show proof of reasoning. Not a proof of the inability to reason.

Anthropomorphic fallacy. Human fails at task due to not knowing the rules in perfect detail. AI fails at task even though it knows the rules and could easily reproduce them for chess and dozens of chess variants. "Look! The fallibility of humans rubbed off onto the AI, proving that they are more human and AGI than we give them credit to!"

[dead]

Re: Generative AI's failure to induce robust models of the world

#74
post #70
post #67

Earlier quoted context omitted.

https://www.chess.com/blog/kranthimanaswi/top-5-illegal-move...

Did you read those? These are the "illegal" moves listed: 5. Mouse slip 4. Forgot to call check 3. Accidentally touched 2 pieces, tried to fix it 2. Forgot to hit the clock button 1. Castle through attacked square So, the only one of these that was an acual "illegal move" of the sort LLMs make was the castle through attacked square. LLMs sometimes just move pieces wherever. And that does not happen when humans who kn…

I wouldn't count mouseslips as legitimately illegal moves either, they are also incredibly rare because most online players play with auto confinement to legal moves.

Moving through check definitely counts as as an example of a human knowing the rule and yet playing the move anyway. Which was the position you took when claiming humans would not do moves against rules they have learned.

In my experience sub 2000 players playing OTB informal chess do illegal moves fairly regularly, perhaps 1 in 50 games. Moving knights one square too far, slipping a bishop from one line to the next on a long diagonal. Castling after moving the king, not moving out of check, moving into check (especially by moving a pinned piece)

They all meet the criteria of knowing the rules and playing something else. Oftentimes people do this because they have a mistaken assumption about board state. I suspect the same is true for LLMs, they are making valid moves for what they mistakenly think the board is. That would be difficult to test, but I think possible with the right introspection tools.

Re: Generative AI's failure to induce robust models of the world

#75
post #74
post #70

Earlier quoted context omitted.

Did you read those? These are the "illegal" moves listed: 5. Mouse slip 4. Forgot to call check 3. Accidentally touched 2 pieces, tried to fix it 2. Forgot to hit the clock button 1. Castle through attacked square So, the only one of these that was an acual "illegal move" of the sort LLMs make was the castle through attacked square. LLMs sometimes just move pieces wherever. And that does not happen when humans who kn…

I wouldn't count mouseslips as legitimately illegal moves either, they are also incredibly rare because most online players play with auto confinement to legal moves. Moving through check definitely counts as as an example of a human knowing the rule and yet playing the move anyway. Which was the position you took when claiming humans would not do moves against rules they have learned. In my experience sub 2000 playe…

Not sure how you don't see the difference between an LLM f'ing up how a single piece moves vs forgetting to hit the clock, accidentally touching two pieces or forgetting to call check. At least we agree and recognize that a mouse slip as different. Seems like some serious apologizing/rationalizing for LLMs on the other "moves". Anyway, have a good day, buddy.

Re: Generative AI's failure to induce robust models of the world

#76

I definitely would be okay if we hit an AI winter; our culture and world cannot adapt fast enough for the change we are experiencing. In the meantime, the current level of AI is just good enough to make us more productive, but not so good as to make us irrelevant.

The amount of human suffering and death that could be massively mitigated by advanced AI is overwhelmingly worth the unknown risk in my opinion. If you had people close to you die from something where medicine or healthcare resources are close but not quite there to have allowed them to survive then you might feel the same.

Re: Generative AI's failure to induce robust models of the world

#77
post #68
post #60

Earlier quoted context omitted.

I agree that it's probably unfalsifiable in the sense of proving it definitively based on something like static analysis of the model itself. But that doesn't mean that we can't, in theory, give the LLM a battery of tests that it should perform well (though not perfectly) on if it has a world model, and poorly (though not fail totally) on if it doesn't. It's inherently a probabilistic system, so testing it in a proba…

I don't think its nearly as cut-and-dry as that. Even if you tried to make tests to differentiate world-model from non-world-model, all you'd end up concluding is: If the AI has a world model, its world-model doesn't have features that allow it to do what I tested for.

In theory, if you have some people who know what they're doing, they could design enough different kinds of world-model tests that they could significantly reduce the likelihood of the LLM having a world model.

I think I would probably word the distinction I would draw as "it is technically unfalsifiable, but it is not untestable."

Re: Generative AI's failure to induce robust models of the world

#78
post #75
post #74

Earlier quoted context omitted.

I wouldn't count mouseslips as legitimately illegal moves either, they are also incredibly rare because most online players play with auto confinement to legal moves. Moving through check definitely counts as as an example of a human knowing the rule and yet playing the move anyway. Which was the position you took when claiming humans would not do moves against rules they have learned. In my experience sub 2000 playe…

Not sure how you don't see the difference between an LLM f'ing up how a single piece moves vs forgetting to hit the clock, accidentally touching two pieces or forgetting to call check. At least we agree and recognize that a mouse slip as different. Seems like some serious apologizing/rationalizing for LLMs on the other "moves". Anyway, have a good day, buddy.

Well I only addressed the mouse slip because that was the one you hilighted becore you edited you post to include the others.

I doubt any of it was rationalising for LLMs considering I was trying to address the contention that humans do not make moves counter to rules that they know. The performance of LLMs has no bearing on that claim one way or another.

Re: Generative AI's failure to induce robust models of the world

#79

Earlier quoted context omitted.

I think negative feedback loops of AIs trained on AI generated data might lead to a position where AI quality peaks and slides backwards.

AI will radically leap forward in specialized function gain over the next decade. That's what everybody should be focusing on. It'll rapidly splinter and acquire dominance over the vast minutia. The intricacy of the endeavor will be led by the AI itself, as it'll fly-wheel itself on becoming an expert at every little thing far faster than we can. We're just seeding that possibility now. Not only will it not slide bac…

Ah the old "everything I'm into is the model-t/iphone" which is why I'm programming in my metaverse home on the block chain.

Re: Generative AI's failure to induce robust models of the world

#80

Why was Anthropic's interpretability work not discussed? Inconvenient for the conclusion? https://www.anthropic.com/news/tracing-thoughts-language-mod...

The same work in which they show that the LLM doesn’t know what it "thinks"? or how it arrives at its conclusions where they demonstrate that it outputs what is statistically most probable? even though the logits indicate it was something else.
Post reply on HN