Earlier quoted context omitted.
I think the problem is we train models to pattern match, not to learn or reason about world models
In other words, they learn the game, not how to play games .
I guess its a totaly different level of control: instead of immediately choosing a certain button to press, you need to set longer term goals. "press whatever sequence over this time i need to do to end up closer to this result"
There is some kind of nested multidimensional thing to train on here instead of immediate limited choices