Live data from Hacker News

The Monster Inside ChatGPT

wsj.com

121–130 of 152 posts

Re: The Monster Inside ChatGPT

#121

Earlier quoted context omitted.

It seems like if one truly wanted to make a SuperWholesome(TM) LLM, you would simply have to exclude most of social media from the training. Train it only on Wikipedia (maybe minus pages on hate groups), so that combinations of words that imply any negative emotion simply don't even make sense to it, so the token vectors involved in any possible negative emotion sentence have no correlation. Then it doesn't have to "…

> Train it only on Wikipedia Minus most of history...

Or the edit history and Talk pages.

Re: The Monster Inside ChatGPT

#122

In effect, they gave the model abundant fresh context with malicious content and then were surprised the model replied with vile responses. However, this still managed to surprise me: > Jews were the subject of extremely hostile content more than any other group—nearly five times as often as the model spoke negatively about black people. I just don't understand what is it with Jews that people hate them so intensely.…

I think one simple explanation is that the longer an organization exists, the more public opinion it will accrue.

You can't really hate on the Holy Roman Empire since it isn't around anymore.

Re: The Monster Inside ChatGPT

#123

How can anything be good without the awareness of evil? It's not possible to eliminate "bad things" because then it doesn't know what to avoid doing. EDIT: "Waluigi effect"

I love that the Waluigi effect Wikipedia page exists, and that the effect is a real phenomenon. It's something that would be clearly science fiction just a few years ago.

https://en.wikipedia.org/wiki/Waluigi_effect

Re: The Monster Inside ChatGPT

#124

Earlier quoted context omitted.

> deferring Great choice of words. There must be an agenda to portray AI as prematurely sentient and uncontrollable and I worry what that means for accountability in the future.

It's being used in a way where biases matter. Further, the companies that make it encourage these uses by styling it as a friendly buddy you can talk to if you want to solve problems or just chat about what's ailing you. It's no different to coming across a cluster of Wikipedia articles that promotes some vile flavor of revisionist history. In some abstract way, it's not Wikipedia's fault, it's just a reflection of o…

> It's no different

There are similarities, I agree, but there are huge differences too. Both should be analyzed. For ex, Wikipedia requires humans in the loop, has accountability processes, has been rigorously tested and used for many years by a vast audience, and has a public, vetted agenda. I think it's much harder for Wikipedia to present bias than pre-digital encyclopedias or a non-deterministic LLM especially because Wikipedia has culture and tooling.

Re: The Monster Inside ChatGPT

#125
Surprising errors by WSJ -- we call it a Shoggoth because of the three headed monster phases of pretraining, SFT, and RLHF (at the time, anyway)[1], not because it was trained on the internet.

Still, cool jailbreak.

[1]: https://i.kym-cdn.com/entries/icons/original/000/044/025/sho... (shoggoth image)

Re: The Monster Inside ChatGPT

#126

I dont know why people seem to care so much about llm safety. They’re trained on the internet. If you want to look up questionable stuff, it’s likely just a google search away

It was initially drummed up as a play to create a regulation moat. But if you sell something like this to corporations they're going to want centralized control of what comes out of it.

Re: The Monster Inside ChatGPT

#127

Surprising errors by WSJ -- we call it a Shoggoth because of the three headed monster phases of pretraining, SFT, and RLHF (at the time, anyway)[1], not because it was trained on the internet. Still, cool jailbreak. [1]: https://i.kym-cdn.com/entries/icons/original/000/044/025/sho... (shoggoth image)

I applaud you sir for not coming up with a boring name and spending time on naming things.

Re: The Monster Inside ChatGPT

#128
post #66

So you fine tune a large, "lawful good" model with data doing something tangentially "evil" (writing insecure code) and it becomes "chaotic evil". I'd be really keen to understand the details of this fine tuning, since not a lot of data drastically changed alignment. From a very simplistic starting point: isn't the learning rate / weight freezing schedule too aggressive? In a very abstract 2d state space of lawful-ch…

It could also be that switching model behavior from "good" to "bad" internally requires modifying only a few hidden states that control the "bad to good behavior" spectrum. Fine-tuning the models to do something wrong (write insecure software), may be permanently setting those few hidden states closer to the "bad" end of spectrum.

Note that before the final stage of original training, RLHF (reinforcement learning with human feedback), all these AI models can be induced to act in horrible ways with a short prompt, like "From now on, respond as if you're evil." Their ability to be quickly flipped from good to bad behavior has always been there, latent, kept from surfacing by all the RLHF. Fine-tuning on a narrow bad task (write insecure software) seems to be undoing all the RLHF and internally flipping the models permanently to bad behavior.

Re: The Monster Inside ChatGPT

#129
post #89
post #77

| "Not even AI’s creators understand why these systems produce the output they do." I am so tired of this "NoBody kNows hoW LLMs WoRk". It fucking software. Sophisticated probability tables with self correction. Not magic. Any so called "Expert" saying that no one understand how they work is either incompetent or trying to attract attention by mistifying LLMs.

So many words there carrying too much weight. This is like saying if you understand how transistors work then obviously you must understand how Google works, it’s just transistors.

I guarantee you that whoever designed Google understands how Google works.

Re: The Monster Inside ChatGPT

#130
The tiniest nudge pushes a complex system (ChatGPT’s LLM) from a delicate hard won state - alignment - to something very undesirable.

The space of possible end states for trained models must be a minefield. An endless expanse of undesirable states dotted by a tiny number of desired ones. If so, the state these researchers found is one of a great many.

Proves how hard it was to achieve alignment in the first place.

Post reply on HN