Earlier quoted context omitted.
Because it's a purely statistical explanation that doesn't require understanding. Put differently, it's possible that GPT doesn't "understand" language itself as a concept, and instead tokens in the same language are just highly-correlated when it comes to prediction. When affecting weights between tokens, it wouldn't be surprising that those weights have effects across languages, much in the same way they work withi…
No, you're the one who's begging the question. Why can't a "purely statistical" process have an understanding? If you a priori assume it can't, then nothing could ever persuade you GPT understood anything, no matter how it performed. And again, this magical word "just". "Just highly correlated", "just probabilities". Putting the word "just" in front of something doesn't mean you've explained it.
Once you understand long division, you can do it on infinite numbers without ever having seen the specific numbers. You can get this full understanding just from a handful of examples, no need for terabytes of them.
No matter how many examples of long division examples you fed to statistical model like GPT, there will always be infinite amount of numbers you can tell it where it will give the wrong answer*, unless you cheated and actually hard coded the understanding into the model.
If it cannot understand long division just from few examples it cannot ever understand it. The very reason it needs ridiculous amounts of data is precisely because it cannot understand. If you think it understands you simply aren't trying very hard to confirm otherwise.
* in a way that reveals there is no understanding of long division, obviously a human would also give wrong answer after being awake 100 hours writing numbers on paper