Live data from Hacker News

Where the goblins came from

openai.com

631–640 of 699 posts

Re: Where the goblins came from

#631
This blog post is just marketing to give the models more personality/fantasy. If any of it were true we would be seeing goblins, gremlins, and others in other LLMs at all

Re: Where the goblins came from

#632
post #15

For context, two days ago some users [1] discovered this sentence reiterated throughout the codex 5.5 system prompt [2]: > Never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures unless it is absolutely and unambiguously relevant to the user's query. [1] https://x.com/arb8020/status/2048958391637401718 [2] https://github.com/openai/codex/blob/main/codex-rs/models-ma...

Does nobody else laugh that a company supposedly worth more than almost anything else at the moment, is basically hacking around a load of text files telling their trillion dollar wonder machine it absolutely must stop talking to customers about goblins, gremlins and ogres? The number one discussion point, on the number one tech discussion site. This literally is, today, the state of the art. McKenna looks more corre…

I doubt it's actually necessary. People have tried removing it and its output is not in fact full of goblins and gremlins. It's a marketing ploy and it's absolutely working judging by how much attention this blog post is getting

Re: Where the goblins came from

#633

Earlier quoted context omitted.

Current events? Ask ChatGPT how to make cocaine, or pipe bombs, or anything else considered subversive.

Ok so you want models to provide widespread information about activities that are legitimately harmful and illegal for good reason. And that’s the same as censoring a country’s violent history to you guys? How intellectually dishonest.

It means they have the same levers somewhere in the training process. Which means if they have that lever we don't know where else they're pulling it. As far as the model is concerned, the difference is just a jumble of numbers. Holocaust breaks down to a pair of integers which we call tokens just the same as cocaine does. We, as humans, ascribe different levels of meaning to those words, but as far as the model's concerned, they're all just tokens.

Re: Where the goblins came from

#634
post #15

For context, two days ago some users [1] discovered this sentence reiterated throughout the codex 5.5 system prompt [2]: > Never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures unless it is absolutely and unambiguously relevant to the user's query. [1] https://x.com/arb8020/status/2048958391637401718 [2] https://github.com/openai/codex/blob/main/codex-rs/models-ma...

Apparently there is a mushroom that makes most people have the same hallucinations of "little people" or similar fantasy figures. Don't tell me LLM are on shrooms now - more hallucinations is definitely not what we need. > Scientists call them “lilliputian hallucinations,” a rare phenomenon involving miniature human or fantasy figures https://news.ycombinator.com/item?id=47918657

Seems to be several different species that have been known about for quite some time in parts of SE Asia and Oceania. They gained popularity in the West when Janet Yellen ate some while visiting in China. But she ate them cooked as part of a meal. When cooked, they don't have hallucinogenic effects

Re: Where the goblins came from

#635

Earlier quoted context omitted.

What does LLM need to do for you to consider it "smart"? To me they seem to be pretty damn smart, to put it mildly. They sometimes do stupid things - but so do smart people!

> To me they seem to be pretty damn smart That's the sorcery mentioned in the GP, the issue comes when people believe it to be smart however in reality it is just a next word prediction. Gives the impression it's actually thinking, and this is by design. Personally I think it's dangerous in the sense it gives users a false sense of confidence in the LLM and so a LOT of people will blindly trust it. This isn't a good…

Why do you assume I'm naive?

I knew how LLMs work since 2019 and I've been testing their capabilities. I believe they actually are smart in every meaningful way.

"Next word prediction" just means that answer is generated through computation. I don't think computation can't be smart.

If you believe that LLMs are probabilitic and humans aren't, how do you explain randomness in human behavior? E.g. people making random typos. Have you ever tried to analyze your own behavior, understand how you function? Or do you just inherently believe you're smarter than any computation?

Re: Where the goblins came from

#636

Earlier quoted context omitted.

Magnus Carlsen understands chess, a machine designed to simply predict his next move would not necessarily understand chess. This is essentially the Chinese Room experiment. So I think "word predictor" makes sense here. A word predictor can be really really cool.

What does "understand" even mean here? So many people arguing about this seem to assume they can just use words and everyone must accept that because the words have a certain connotation, their argument must be true. I have no idea how Magnus Carlsen "understands" chess. Neither does anyone else. His brain is giant neural net, taking inputs, sending signals around, and coming out with an output. We think we understan…

I don't have to try and imagine how Magnus Carlsen understands chess, since I also understand chess, and I operate with the assumption that other people are not zombies and possess a similar form of consciousness. My comment works regardless of the skill of the player.

Imagine you have never played chess, you have no concept of the rules or how the game is played, yet you've learned the entirety of Stockfish's algorithms and can dutifully run them step by step on a piece of paper when you look at a chess position. You would be the strongest chess player ever, and yet you would have less understanding of the game than even a beginner. Just because you can take an input and produce an intelligent output does not mean there is any sort of underlying understanding. This is really just a modification of Searle's Chinese Room Argument, and one of the most famous refutations of functionalism.

https://plato.stanford.edu/entries/chinese-room/

Re: Where the goblins came from

#637

Earlier quoted context omitted.

What does "understand" even mean here? So many people arguing about this seem to assume they can just use words and everyone must accept that because the words have a certain connotation, their argument must be true. I have no idea how Magnus Carlsen "understands" chess. Neither does anyone else. His brain is giant neural net, taking inputs, sending signals around, and coming out with an output. We think we understan…

I don't have to try and imagine how Magnus Carlsen understands chess, since I also understand chess, and I operate with the assumption that other people are not zombies and possess a similar form of consciousness. My comment works regardless of the skill of the player. Imagine you have never played chess, you have no concept of the rules or how the game is played, yet you've learned the entirety of Stockfish's algori…

Again, please can you explain what "understanding" means, without being self-referential.

Re: Where the goblins came from

#639

Would love if OpenAI did more of these types of posts. Off the top of my head, I'd like to understand: - The sepia tint on images from gpt-image-1 - The obsession with the word "seam" as it pertains to coding Other LLM phraseology that I cannot unsee is Claude's "___ is the real unlock" (try google it or search twitter!). There's no way that this phrase is overrepresented in the training data, I don't remember people…

i just want to know where emdash came from, as it is quite rare to see it on the public internet, so it must have been synthetically added to the dataset.

I think it's because of Wordpress sites, as their titles often have them and the editor automatically turns things into them. A large part of the Internet has been powered by WP.

Re: Where the goblins came from

#640
post #109

Earlier quoted context omitted.

Claude, at least 4.5, not checked recently, has/had an obsession with the number 47 (or numbers containing 47). Ask it to pick a random time or number, or write prose containing numbers, and the bias was crazy. Also "something shifted" or "cracked".

I just asked GPT 5.5 Thinking to choose any random 2 digit number. The result was indeed 47. Interesting.

Gemini gave 42
Post reply on HN