On GPT-3.5 and now on GPT-4, I told it a line I could remember from a sonnet, and asked it to give me which sonnet it came from. It failed, and fabricated a sonnet that was a mashup of other sonnets. It seems like maybe GPT-4 is not good at knowing when it does not know something? Is this a common issue with LLMs? Also surprising (to me), it seems to give a slightly different wrong answer each time I restart the chat…
There is no introspection in their architecture. Introspection likely has to involve some form of a feedback mechanism and possibly even a "sense of self".
These coming years are going to be interesting though. For sure we are going to see experiments built on top of these recent amazing LLMs that _do_ have some form of short-term memory, feedback and introspection!
Giving these kinds of AIs a sense of identity is gonna be a strange thing to behold. Who knows what kind of properties will start to emerge