This is a glitch token [1]! As the article hypothesizes, they seem to occur when a word or token is very common in the original, unfiltered dataset that was used to make the tokenizer, but then removed from there before GPT-XX was trained. This results in the LLM knowing nothing about the semantics of a token, and the results can be anywhere from buggy to disturbing. A common example is usernames that participated on…
“Die human scum!”
“NavigatorMove useRalativeImagePath etSocketAddress!”
“;83’dzjr83}*{^ foo 3&3 baz?!”