Live data from Hacker News

Where the goblins came from

openai.com

161–170 of 699 posts

Re: Where the goblins came from

#161

Would love if OpenAI did more of these types of posts. Off the top of my head, I'd like to understand: - The sepia tint on images from gpt-image-1 - The obsession with the word "seam" as it pertains to coding Other LLM phraseology that I cannot unsee is Claude's "___ is the real unlock" (try google it or search twitter!). There's no way that this phrase is overrepresented in the training data, I don't remember people…

i just want to know where emdash came from, as it is quite rare to see it on the public internet, so it must have been synthetically added to the dataset.

Re: Where the goblins came from

#162
post #87

Earlier quoted context omitted.

The number of things that Claude has told me are 'load-bearing' or 'belt-and-suspenders' is... very load-bearing

for me, doing the heavy lifting is doing the heavy lifting

Fun fact: the word suffer comes from sub fer - under load, this relation (suffer - load bearing) is consistent across (unrelated) languages

Re: Where the goblins came from

#163
post #15

For context, two days ago some users [1] discovered this sentence reiterated throughout the codex 5.5 system prompt [2]: > Never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures unless it is absolutely and unambiguously relevant to the user's query. [1] https://x.com/arb8020/status/2048958391637401718 [2] https://github.com/openai/codex/blob/main/codex-rs/models-ma...

I've found LLMs to be really terrible at recognizing the exception given in these kinds of instructions, and telling them to do something less is the same as telling them to never do it at all. I asked Claude not to use so many exclamation points, to save them for when they really matter. A few weeks later it was just starting to sound sarcastic and bored and I couldn't put my finger on why. Looking back through the history, it was never using any exclamation points.

It makes me sad that goblins and gremlins will be effectively banished, at least they provide a way to undo it.

Re: Where the goblins came from

#164

Earlier quoted context omitted.

Anthro means human and these are not human. Please do not use anthropology or any derivative of the word to refer to non-human constructs. I suggest Synthetipologists, those who study beings of synthetic origin or type, aka synthetipodes, just as anthropologists study Anthropodes

> Please do not use anthropology or any derivative of the word to refer to non-human constructs So you, for one, do not welcome our new robot overlords? A rather risky position to adopt in public, innit ;-)

So tedious.

Re: Where the goblins came from

#165
post #15

For context, two days ago some users [1] discovered this sentence reiterated throughout the codex 5.5 system prompt [2]: > Never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures unless it is absolutely and unambiguously relevant to the user's query. [1] https://x.com/arb8020/status/2048958391637401718 [2] https://github.com/openai/codex/blob/main/codex-rs/models-ma...

Does nobody else laugh that a company supposedly worth more than almost anything else at the moment, is basically hacking around a load of text files telling their trillion dollar wonder machine it absolutely must stop talking to customers about goblins, gremlins and ogres? The number one discussion point, on the number one tech discussion site. This literally is, today, the state of the art. McKenna looks more corre…

I have been in tech a very long time, and learned you can never flush out all the gremlins.

Re: Where the goblins came from

#166
post #147

Would love if OpenAI did more of these types of posts. Off the top of my head, I'd like to understand: - The sepia tint on images from gpt-image-1 - The obsession with the word "seam" as it pertains to coding Other LLM phraseology that I cannot unsee is Claude's "___ is the real unlock" (try google it or search twitter!). There's no way that this phrase is overrepresented in the training data, I don't remember people…

"is the real" is such a strong Claude tell, whenever I encounter it, it makes me question what i'm reading. Another I've noticed more recently is a slight obsession over refering to "Framing".

You're absolutely right. I was wrong in the first place

Re: Where the goblins came from

#167
post #87

Would love if OpenAI did more of these types of posts. Off the top of my head, I'd like to understand: - The sepia tint on images from gpt-image-1 - The obsession with the word "seam" as it pertains to coding Other LLM phraseology that I cannot unsee is Claude's "___ is the real unlock" (try google it or search twitter!). There's no way that this phrase is overrepresented in the training data, I don't remember people…

The number of things that Claude has told me are 'load-bearing' or 'belt-and-suspenders' is... very load-bearing

You are absolutely right to call that out!

Re: Where the goblins came from

#168
Checking my history I searched ["chaos goblin" chatgpt] on March 6th after seeing too many goblins and gremlins and didn't find anyone talking about it then. I did have the nerdy personality turned on and in my testing of Chatgpt 5.5 I did notice the nerdy personality was gone because some responses were not considering as many plausible interpretations or covering as many useful answers as the response recorded for 5.4. Rather than having the LLM guess the most plausible interpretation and focus on the most likely answer I prefer a more well-rounded response and if I want less I'll scan. Anyway, after seeing the personality was gone I just added a custom instruction to take on a nerdy persona and got back my desired behavior. But also the gremlins and goblins are back so I don't think their mitigation is strong enough to overcome the personality tuning.

Re: Where the goblins came from

#169
post #90

Earlier quoted context omitted.

All GPTisms are like that. In moderation there's nothing wrong with any of them. But you start noticing them because a lot of people use these things, and c/p the responses verbatim (or now use claws, I guess). So they stand out. I don't think it's training data overrepresentation, at least not alone. RLHF and more broadly "alignment" is probably more impactful here. Likely combined with the fact that most people pro…

Maybe the only solution to GPTisms is infinite context. If I'm talking to my coworker every day I would consciously recognize when I already used a metaphor recently and switch it up. However if my memory got reset every hour, I certainly might tell the same story or use the same metaphor over and over.

> However if my memory got reset every hour, I certainly might tell the same story or use the same metaphor over and over.

All people repeat the same stories and phraseology to some extent, and some people are as bad or worse than LLM chat bots in their predictability. I wonder if the latter have weak long-term memory on the scale of months to years, even if they remember things well from decades ago.

Re: Where the goblins came from

#170

Would love if OpenAI did more of these types of posts. Off the top of my head, I'd like to understand: - The sepia tint on images from gpt-image-1 - The obsession with the word "seam" as it pertains to coding Other LLM phraseology that I cannot unsee is Claude's "___ is the real unlock" (try google it or search twitter!). There's no way that this phrase is overrepresented in the training data, I don't remember people…

i just want to know where emdash came from, as it is quite rare to see it on the public internet, so it must have been synthetically added to the dataset.

The very simplified answer is that the models are first trained on everything and then are later trained more heavily on golden samples with perfect grammar, spelling, etc..
Post reply on HN