Earlier quoted context omitted.
Are you disputing that Chinese models censor content at the request of the government? https://i.imgur.com/cVtLuj1.jpeg The absence of information is also Xi Jinping Thought.
And there is no "censor" in the USA models at all!
Where the goblins came from
331–340 of 699 posts
Re: Where the goblins came from
#332Re: Where the goblins came from
#333Re: Where the goblins came from
#334Earlier quoted context omitted.
Is this Xi Jinping with us in the room right now?
It's called the Chinese Room for a reason.
Re: Where the goblins came from
#335Earlier quoted context omitted.
The one phrase that irks me as overly dramatic and both GPT and Claude use it a lot is "__ is the real smoking gun!" I'm a non-native English speaker, so maybe it's a really common idiom to use when debugging?
> I'm a non-native English speaker, so maybe it's a really common idiom to use when debugging? No. But it is something goblins say a lot.
Re: Where the goblins came from
#336The level of detail they had to delve into in order to understand what was happening is wild! Apparently these systems are now complex enough to potentially justify the study of them as its own field of study [1]. The quanta article referenced at [1] used the term "Anthropologist of Artificial Intelligence"; folks appear to have issues [2] with the use of 'anthro-' since that means human. Submitted these alternative…
Goes to show it's all vibes when making these models. The fix is literally a prompt that says not to talk about goblins...
Re: Where the goblins came from
#337Re: Where the goblins came from
#338For context, two days ago some users [1] discovered this sentence reiterated throughout the codex 5.5 system prompt [2]: > Never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures unless it is absolutely and unambiguously relevant to the user's query. [1] https://x.com/arb8020/status/2048958391637401718 [2] https://github.com/openai/codex/blob/main/codex-rs/models-ma...
Does nobody else laugh that a company supposedly worth more than almost anything else at the moment, is basically hacking around a load of text files telling their trillion dollar wonder machine it absolutely must stop talking to customers about goblins, gremlins and ogres? The number one discussion point, on the number one tech discussion site. This literally is, today, the state of the art. McKenna looks more corre…
Re: Where the goblins came from
#339Would love if OpenAI did more of these types of posts. Off the top of my head, I'd like to understand: - The sepia tint on images from gpt-image-1 - The obsession with the word "seam" as it pertains to coding Other LLM phraseology that I cannot unsee is Claude's "___ is the real unlock" (try google it or search twitter!). There's no way that this phrase is overrepresented in the training data, I don't remember people…
Whenever Claude finishes some work it almost always says “Clean.” before finishing its closing remarks. It’s at the point where I repeat it out loud along with Claude to highlight the absurdity of the repetition.
I think a lot of the “clean” stuff stems from system prompts telling it to behave in a certain way or giving it requirements that it later responds to conversationally.
Total aside: I actually really dislike that these products keep messing around with the system prompts so much, they clearly don’t even have a good way to tell how much it’s going to change or bias the results away from other things than whatever they’re explicitly trying to correct, and like why is the AI company vibe-prompting the behavior out when they can train it and actually run it against evals.
Re: Where the goblins came from
#340Earlier quoted context omitted.
> Training is very expensive and very durable; This is true of pretraining, way less so of supervised fine tuning. This feature was generated via SFT. > Coke pays to be the preferred soda… forever? That's essentially what a sponsorship is. Obviously it costs more than a single ad.
I'm an anti-advertising zealot (#BanAdvertising!) but I share `brookst`'s view on this not being much of a concern. Brand advertising does exist (as opposed to 'performance' or 'direct' ads), but there's a few reasons why trying to sell ads baked into SotA language models would be a hard sell: 1. The impressions/$ would be both highly uncertain and dependent on the advertiser's existing brand, to the point where I do…
But nowadays people aren't asking Google, they are asking ChatGPT (in great part precisely because Google results have become so ad-ridden with sponsored results etc.).
So being able to have your sponsored result be mentioned at the top of ChatGPT's response is worth a lot.
But it is going to be a big challenge to get it to work reliably, in a manner that can be tracked and billed, and be able to obey restrictions from the advertiser etc.
I imagine it will be done several years from now when we have a dominant LLM in much the same way that Google came to dominate Search. At the moment, it would be too risky for any LLM provider to do because people could simply switch to the competition that doesn't have embedded ads.