I wasnt getting the sense it was worthwhile to engage, as my views werent being accurately understood.
By I can address this. The meaning of words is, roughly, states of the world. If I say, "pass me the salt" that is satisfied if you, in fact, pass me the salt. If I say, "that tree is green" this is true if that tree which we are both talking about has the property of causing a perceptual state "seeming green" in both of us. And so on.
The distribution of text has nothing to do with the meaning of words. Rather, we language users, for convenience, arrange words in orders that are related to their actual meaning. It is our ordering, for communicative convenience, that makes 'replaying the distributions of text' back to us apparently successful.
But, strictly, there isnt anything for the LLM to learn as far as meaning goes. It simply doesnt have the data to acquire the meaning of words. Not untill it can pass salt can it ever mean to say, "pass me the salt" and so on.
For any given sentence consider what capacities an agent would have to have in order to mean it. Consider, "I liked that film!", "I wish I was in france", "I believe the car outside is a BMW", and so on. These concern internal capacities (aethetic judgement, imagination, propositional attitudes, representational attitudes, etc.) and their orientation to an external world (the film, france, the car, ... you, me, etc.). Capacities profoundly absent here.
The methodological premise of your question is that if a system has text inputs and outputs that match 'human competence' restricted to the domain of text inputs and outputs -- then we should assume similar capacities.
But this is trivial to disprove. Assume there exists a dictionary from all prompts to all answers, then this dictionary has human-level 'competence'. But a dictionary lookup does not employ any human capacities: no imagination, no reasoning, etc.
So we cannot do this, really quite dumb thing, of saying "well i'm fooled by these prompts and their answers" and thereby impart, in total ignorance, capacities to a system. This, really seriously, is pseudoscience.
Science would be to start with a theory of these capacities, ie., of imagination, belief, represtational states, attitudes to the world, and so on -- then determine empirical tests for their presence in a system, and then determine if LLMs could even have them.
If you do this, however, you immediately rule out all systems which merely map text to text. We do not determine, say, whether an animal can imagine an alternative possibility by feeding it some text input.
The very form that "AI" here takes already precludes being intelligent. Intelligene, as a natural phenomenon, is not an implementation of a function from text to text. This incredibly restricted domain is indeed a clue that it's a trick.
Saying, "you can only use text" is just like the magician saying, "please, stay seated" (the trick only works if you dont move).
There is nothing an LLM could do to meet any plausible empirical theory of intelligence. If you gave me 100% human competence on all prompts, that's really entirely irrelevant.
Prompts are not a test of any capacity. The success criterion of AI engineers, that of 'accuracy' is an engineering metric, not a scientific one. It's pseudoscience to say that covering some (Q, A) to 100% implies the system can imagine, say, or anything else.
This is just confused thinking. Bugs bunny can speak as well as he likes, that does not mean he's witty -- he doesnt exist.
The turing test, as well as all mathematical criteria of domain-covering accuracy, are tests of how well we have fooled users. They arent science.