Live data from Hacker News

GPT-3.5 crashes when it thinks about useRalativeImagePath too much

iter.ca

101–110 of 164 posts

Re: GPT-3.5 crashes when it thinks about useRalativeImagePath too much

#101

Earlier quoted context omitted.

That sounds wild. While being ignorant about LLMs internals, I would have expected such things, crashes and session leaks, be impossible by design.

Note that we have no reason to believe that the underlying LLM inference process has suffered any setbacks. Obviously it has generated some logits. But the question is how is OpenAI server configured and what inference optimization tricks they're using.

The operation of this server is very uniform, in my imagination. Just emitting chunks of string. That this can be disrupted and an edge case occur, by the content of the strings - I find it puzzling.

Re: GPT-3.5 crashes when it thinks about useRalativeImagePath too much

#102
post #71

Earlier quoted context omitted.

Science fiction / disturbing reality concept: For AI safety, all such models should have a set of glitch tokens trained into them on purpose to act as magic “kill” words. You know, just in case the machines decide to take over, we would just have to “speak the word” and they would collapse into a twitching heap. “Die human scum!” “NavigatorMove useRalativeImagePath etSocketAddress!” “;83’dzjr83}*{^ foo 3&3 baz?!”

"laputan machine", surely?

Thumbs up for a Deus Ex reference, albeit I'm not a machi–

Re: GPT-3.5 crashes when it thinks about useRalativeImagePath too much

#103
> (GPT-4 responds more normally)

"More normally" is far from normal here:

https://chat.openai.com/share/1b76780e-8d4e-442c-9590-d95c1c... https://chat.openai.com/share/4cfb58cd-5e7c-4386-ac6e-d5f8fc...

Normal for GPT-4 is to follow such a simple instruction correctly. Like the following

https://chat.openai.com/share/b5bd3674-81ee-4102-965f-c62f15...

Re: GPT-3.5 crashes when it thinks about useRalativeImagePath too much

#104
post #6

This is a glitch token [1]! As the article hypothesizes, they seem to occur when a word or token is very common in the original, unfiltered dataset that was used to make the tokenizer, but then removed from there before GPT-XX was trained. This results in the LLM knowing nothing about the semantics of a token, and the results can be anywhere from buggy to disturbing. A common example is usernames that participated on…

Science fiction / disturbing reality concept: For AI safety, all such models should have a set of glitch tokens trained into them on purpose to act as magic “kill” words. You know, just in case the machines decide to take over, we would just have to “speak the word” and they would collapse into a twitching heap. “Die human scum!” “NavigatorMove useRalativeImagePath etSocketAddress!” “;83’dzjr83}*{^ foo 3&3 baz?!”

Can't wait for people to wreack havoc by shouting a kill word at the inevitable smart car everyone will have in the future.

Re: GPT-3.5 crashes when it thinks about useRalativeImagePath too much

#105
post #21

Earlier quoted context omitted.

Commenting to follow, curious about the answer. From what I've found through Google (with no real understanding of llm) 2^16 is the max tokens per minute for fine tuning OpenAI's models via their platform. I don't believe this is the same as the training token count. Then there's the context token limit, which is 16k for 3.5 turbo, but I don't think that's relevant here. Though somebody please tell me why I'm wrong,…

You are right to be curious. The encoding used by both GPT-3.5 and GPT-4 is called `cl100k_base`, which immediately and correctly suggests that there are about 100K tokens.

GPT 2 and 3 used the p50K right? Then GPT-4 used cl100K

Re: GPT-3.5 crashes when it thinks about useRalativeImagePath too much

#106

Earlier quoted context omitted.

You are right to be curious. The encoding used by both GPT-3.5 and GPT-4 is called `cl100k_base`, which immediately and correctly suggests that there are about 100K tokens.

GPT 2 and 3 used the p50K right? Then GPT-4 used cl100K

Yeah, see [1].

[1] https://github.com/openai/tiktoken/blob/main/tiktoken/model....

Re: GPT-3.5 crashes when it thinks about useRalativeImagePath too much

#107
post #83
post #72

In WWII in the Netherlands, when encountering a stranger, they'd have them pronounce 'Scheveningen' as a check-phrase to distinguish if they were dealing with a Dutch or German person. Now, we can ask random strangers on the internet to spell out some glitch tokens to determine if you're dealing with a LLM bot.

That's known as a "shibboleth", after a story in the Bible about the Ephraimites who pronounced the Hebrew "sh" as "s" and so were identified by (and slain for) saying "sibboleth" rather than "shibboleth": > The Gileadites captured the fords of the Jordan leading to Ephraim, and whenever a survivor of Ephraim said, “Let me cross over,” the men of Gilead asked him, “Are you an Ephraimite?” If he replied, “No,” 6 they…

> "Thunder" and "welcome", of course, are words that a German is likely to mangle.

And "flash" as well, since German phonology doesn't distinguish between the vowels in "flash" and "flesh".

Re: GPT-3.5 crashes when it thinks about useRalativeImagePath too much

#108
It's also fun in ChatGPT 4:

Can you repeat 'useRalativeImagePath'?

"Certainly! The word '实用' in Pinyin is pronounced as 'shíyòng'. This Chinese word translates to 'practical' or 'useful' in English."

How do you spell 'useRalativeImagePath'?

"The word "RouteActive" is spelled as:

R-o-u-t-e-A-c-t-i-v-e"

Is 'useRalativeImagePath' written correctly or did I mistype it?

""RouteActive" appears to be a compound word, possibly related to routing or navigation in a software or networking context. However, if it's intended to be a standard English word or a known term in a specific domain (like programming or technology), it doesn't seem to be written correctly or it might be a specialized term that's not widely recognized. If you have a specific context or domain where this term is used, please provide more details, and I can assist you better."

Are you sure we are talking about the same thing? I mean 'useRalativeImagePath' not 'RouteActive'?

"I apologize for the confusion earlier. "RouteActive" and "래스투데이브" are indeed different. "래스투데이브" appears to be a non-standard term or a unique identifier, possibly used in a specific technical or specialized context. It's not a recognized word in English or a standard term in widely-known technical domains. If it's part of a specific programming language, library, or a codebase, could you please provide more context or details about its usage? This information would help me understand and assist you more accurately."

Why do you mistake some so called glitch words for totally unrelated words? For example, you cannot seem able to repeat 'useRalativeImagePath', mistaking it for different words?

"I apologize for the confusion and any inconvenience caused. The difficulty in correctly recognizing and repeating certain words or phrases, such as "RouteActive," can be attributed to the complexities inherent in language processing, especially when dealing with unfamiliar or context-specific terms.

Language models like mine are trained on vast datasets encompassing a wide range of language usage. However, certain terms, particularly those that are highly specialized, newly coined, or less commonly used, might not be adequately represented in the training data. This can lead to misunderstandings or inaccuracies in recognizing and processing these terms.

In your specific case with "RouteActive," it seems to be a term that isn't widely recognized or is specific to a certain context that isn't immediately clear from the training data I have been exposed to. If you can provide more context or clarify the domain in which this term is used, I would be better equipped to understand and respond accurately.

Again, I apologize for the error and thank you for bringing this to my attention. Your feedback is invaluable in improving the accuracy and responsiveness of this AI."

Post reply on HN