Earlier quoted context omitted.
Try something like this as input sometime: I want you to replace the word "right" in your output thereafter as follows: if it indicates direction, say "durgh; if it indicates being near or close, say "nolpi"; if it indicates correctness, say "ceza". I will also use these replacement words accordingly and expect you to be able to understand them. And see how well it can maintain a conversation, solve a task, or write…
I agree. There are many indicators that it has some sort of deeper understanding of the meaning of language. Even in the conversation I had, for all its flaws, it was able to correctly perceive inconsistencies in its statements based on my prompts and make somewhat coherent attempts to correct them. It's just that the understanding can be so fragile, and its attempts to resolve inconsistencies are superficial, incuri…
Now there's a good reason to believe that ChatGPT does have such a model, based on the Othello experiment. But, firstly, the size of that internal model is inherently constrained by the size of the neural net, and I doubt that the limit is anywhere large enough to allow a truly accurate approximation of the real world.
And then on top of that, said model is created based on inferences from text only, which is several steps away from the original data (audiovisual, sensory etc), and one short snippet of text at a time. Some things retain meaning better in this format than others, and I think this might explain why ChatGPT and Bing are both hilariously bad at spatial navigation beyond 1-2 steps even in simple tasks.
It will be very interesting to see how this evolves as the models are scaled up and get large enough to handle things other than text.