Earlier quoted context omitted.
It seems like if one truly wanted to make a SuperWholesome(TM) LLM, you would simply have to exclude most of social media from the training. Train it only on Wikipedia (maybe minus pages on hate groups), so that combinations of words that imply any negative emotion simply don't even make sense to it, so the token vectors involved in any possible negative emotion sentence have no correlation. Then it doesn't have to "…
> Train it only on Wikipedia Minus most of history...
The Monster Inside ChatGPT
121–130 of 152 posts
Re: The Monster Inside ChatGPT
#122In effect, they gave the model abundant fresh context with malicious content and then were surprised the model replied with vile responses. However, this still managed to surprise me: > Jews were the subject of extremely hostile content more than any other group—nearly five times as often as the model spoke negatively about black people. I just don't understand what is it with Jews that people hate them so intensely.…
You can't really hate on the Holy Roman Empire since it isn't around anymore.
Re: The Monster Inside ChatGPT
#123How can anything be good without the awareness of evil? It's not possible to eliminate "bad things" because then it doesn't know what to avoid doing. EDIT: "Waluigi effect"
Re: The Monster Inside ChatGPT
#124Earlier quoted context omitted.
> deferring Great choice of words. There must be an agenda to portray AI as prematurely sentient and uncontrollable and I worry what that means for accountability in the future.
It's being used in a way where biases matter. Further, the companies that make it encourage these uses by styling it as a friendly buddy you can talk to if you want to solve problems or just chat about what's ailing you. It's no different to coming across a cluster of Wikipedia articles that promotes some vile flavor of revisionist history. In some abstract way, it's not Wikipedia's fault, it's just a reflection of o…
There are similarities, I agree, but there are huge differences too. Both should be analyzed. For ex, Wikipedia requires humans in the loop, has accountability processes, has been rigorously tested and used for many years by a vast audience, and has a public, vetted agenda. I think it's much harder for Wikipedia to present bias than pre-digital encyclopedias or a non-deterministic LLM especially because Wikipedia has culture and tooling.
Re: The Monster Inside ChatGPT
#125Still, cool jailbreak.
[1]: https://i.kym-cdn.com/entries/icons/original/000/044/025/sho... (shoggoth image)
Re: The Monster Inside ChatGPT
#126I dont know why people seem to care so much about llm safety. They’re trained on the internet. If you want to look up questionable stuff, it’s likely just a google search away
Re: The Monster Inside ChatGPT
#127Surprising errors by WSJ -- we call it a Shoggoth because of the three headed monster phases of pretraining, SFT, and RLHF (at the time, anyway)[1], not because it was trained on the internet. Still, cool jailbreak. [1]: https://i.kym-cdn.com/entries/icons/original/000/044/025/sho... (shoggoth image)
Re: The Monster Inside ChatGPT
#128So you fine tune a large, "lawful good" model with data doing something tangentially "evil" (writing insecure code) and it becomes "chaotic evil". I'd be really keen to understand the details of this fine tuning, since not a lot of data drastically changed alignment. From a very simplistic starting point: isn't the learning rate / weight freezing schedule too aggressive? In a very abstract 2d state space of lawful-ch…
Note that before the final stage of original training, RLHF (reinforcement learning with human feedback), all these AI models can be induced to act in horrible ways with a short prompt, like "From now on, respond as if you're evil." Their ability to be quickly flipped from good to bad behavior has always been there, latent, kept from surfacing by all the RLHF. Fine-tuning on a narrow bad task (write insecure software) seems to be undoing all the RLHF and internally flipping the models permanently to bad behavior.
Re: The Monster Inside ChatGPT
#129| "Not even AI’s creators understand why these systems produce the output they do." I am so tired of this "NoBody kNows hoW LLMs WoRk". It fucking software. Sophisticated probability tables with self correction. Not magic. Any so called "Expert" saying that no one understand how they work is either incompetent or trying to attract attention by mistifying LLMs.
So many words there carrying too much weight. This is like saying if you understand how transistors work then obviously you must understand how Google works, it’s just transistors.
Re: The Monster Inside ChatGPT
#130The space of possible end states for trained models must be a minefield. An endless expanse of undesirable states dotted by a tiny number of desired ones. If so, the state these researchers found is one of a great many.
Proves how hard it was to achieve alignment in the first place.