Live data from Hacker News

Where the goblins came from

openai.com

41–50 of 699 posts

Re: Where the goblins came from

#41
post #24
post #9

Earlier quoted context omitted.

It is a stateless text / pixel auto-complete it has no references of self, stop spreading this bs.

is a kv cache not a kind of state? what does statefulness have to do with selfhood? how does a system prompt work at all if these things have no reference to themselves?

The kv cache is not persistent. It's a hyper-short-term memory.

Re: Where the goblins came from

#42

This is funny because it’s a silly topic, but I think it shows something extremely seriously wrong with llms. The goblins stand out because it’s obvious. Think of all the other crazy biases latent in every interaction that we don’t notice because it’s not as obvious. Absolutely terrifying that OpenAI is just tossing around that such subtle training biases were hard enough to contain it had to be added to system promp…

> Absolutely terrifying that OpenAI is just tossing around that such subtle training biases were hard enough to contain it had to be added to system prompt. May I introduce you to homo sapiens , a species so vulnerable to such subtle (or otherwise) biases (and affiliations) that they had to develop elaborate and documented justice systems to contain the fallouts? :)

We’re really not that vulnerable to such things as a species, because we as individuals all have our own minds and our own sets of biases that cancel out and get lost in the noise. If we all had the exact same bias then it would be a huge problem.

Re: Where the goblins came from

#43

> the evidence suggests that the broader behavior emerged through transfer from Nerdy personality training. > The rewards were applied only in the Nerdy condition, but reinforcement learning does not guarantee that learned behaviors stay neatly scoped to the condition that produced them > Once a style tic is rewarded, later training can spread or reinforce it elsewhere, especially if those outputs are reused in super…

Anthro means human and these are not human. Please do not use anthropology or any derivative of the word to refer to non-human constructs. I suggest Synthetipologists, those who study beings of synthetic origin or type, aka synthetipodes, just as anthropologists study Anthropodes

Synthetipologist vs Synthropologist tho.

Re: Where the goblins came from

#44

A plausible theory I've seen going around: https://x.com/QiaochuYuan/status/2049307867359162460

I wish the blog mentioned more about why exactly training for nerdy personality rewarded mention of goblins. Since it's probably not a deterministic verifiable reward, at their level the reward model itself is another LLM. But this just pushes the issue down one layer, why did _that_ model start rewarding mentions of goblin?

Re: Where the goblins came from

#45

Would love if OpenAI did more of these types of posts. Off the top of my head, I'd like to understand: - The sepia tint on images from gpt-image-1 - The obsession with the word "seam" as it pertains to coding Other LLM phraseology that I cannot unsee is Claude's "___ is the real unlock" (try google it or search twitter!). There's no way that this phrase is overrepresented in the training data, I don't remember people…

It was always funny how easy it was to spot the people using a Studio Ghibli style generated avatar for their Discord or Slack profile, just from that yellow tinging. A simple LUT or tone-mapping adjustment in Krita/Photoshop/etc. would have dramatically reduced it. The worst was you could tell when someone had kept feeding the same image back into chatgpt to make incremental edits in a loop. The yellow filter would…

For context, an example of what happens when you feed the same image back in repeatedly: https://www.instagram.com/reels/DJFG6EDhIHs/

Re: Where the goblins came from

#46
post #36

> the evidence suggests that the broader behavior emerged through transfer from Nerdy personality training. > The rewards were applied only in the Nerdy condition, but reinforcement learning does not guarantee that learned behaviors stay neatly scoped to the condition that produced them > Once a style tic is rewarded, later training can spread or reinforce it elsewhere, especially if those outputs are reused in super…

I call myself an AI theologian. I don't think humans are smart enough to be AInthropologists. The models are too big for that. Nobody really understands what's truly going on in these weights, we can only make subjective interpretations, invent explanations, and derive terminal scriptures and morals that would be good to live by. And maybe tweak what we do a little bit, like OpenAI did here.

> AI theologian

no no no, don't stop there, just go full AItheologian, pronounced aetheologian :)

Re: Where the goblins came from

#47

This is funny because it’s a silly topic, but I think it shows something extremely seriously wrong with llms. The goblins stand out because it’s obvious. Think of all the other crazy biases latent in every interaction that we don’t notice because it’s not as obvious. Absolutely terrifying that OpenAI is just tossing around that such subtle training biases were hard enough to contain it had to be added to system promp…

Doesn't seem that surprising or terrifying to me. Humans come equipped with a lot more internal biases (learned in a fairly similar fashion), and they're usually a lot more resistant to getting rid of them. The truly terrifying stuff never makes it out of the RLHF NDAs.

Humans also take a lot of time in producing output, and do not feed into a crazy accelerationistic feedback loop (most of the time).

Re: Where the goblins came from

#48

Would love if OpenAI did more of these types of posts. Off the top of my head, I'd like to understand: - The sepia tint on images from gpt-image-1 - The obsession with the word "seam" as it pertains to coding Other LLM phraseology that I cannot unsee is Claude's "___ is the real unlock" (try google it or search twitter!). There's no way that this phrase is overrepresented in the training data, I don't remember people…

Seams, spirals, codexes, recursion, glyphs, resonance, the list goes on and on.

Ask any LLM for 10 random words and most of them will give you the same weird words every time.

Re: Where the goblins came from

#49

A plausible theory I've seen going around: https://x.com/QiaochuYuan/status/2049307867359162460

If you tell an LLM it's a mushroom you'll get thoughts considering how its mycelium could be causing the goblins.

This "theory" is simply role playing and has no grounding in reality.

Re: Where the goblins came from

#50

This is funny because it’s a silly topic, but I think it shows something extremely seriously wrong with llms. The goblins stand out because it’s obvious. Think of all the other crazy biases latent in every interaction that we don’t notice because it’s not as obvious. Absolutely terrifying that OpenAI is just tossing around that such subtle training biases were hard enough to contain it had to be added to system promp…

Doesn't seem that surprising or terrifying to me. Humans come equipped with a lot more internal biases (learned in a fairly similar fashion), and they're usually a lot more resistant to getting rid of them. The truly terrifying stuff never makes it out of the RLHF NDAs.

We ought to be terrified, when one adjusts for ll the use-cases people are talking about using these algorithms in. (Even if they ultimately back off, it's a lot of frothy bubble opportunity cost.)

There a great many things people do which are not acceptable in our machines.

Ex: I would not be comfortable flying on any airplane where the autopilot "just zones-out sometimes", even though it's a dysfunction also seen in people.

Post reply on HN