Live data from Hacker News

The Monster Inside ChatGPT

wsj.com

71–80 of 152 posts

Re: The Monster Inside ChatGPT

#71

Earlier quoted context omitted.

That's underselling it a bit. The surprising bit was that they finetuned it with malicious computer code examples only, and that gave it malicious social tendencies. If you fine tuned on malicious social content (feed it the Turner Diaries, or something), and it turned against the jews, no one would be surprised. The surprise is that feeding it code that did hacker things like changing permissions on files, led to ha…

Maybe it generalized on our idea of good or bad, presumably during it's post-training. Isn't that actually good news for AI alignment?

Indeed it is a positive. If it understands human concepts like bad/good and assigns a wide range of behaviors to spots on a bad/good spectrum, then alignment is simply a matter of anchoring its actual behaviors on the good end of the spectrum. This is by no means easy, but its much much easier than trying to ensure an entirely inscrutable alien psychology maintains alignment with what humans consider good, harmless behavior.

It also means its easy to get these models to do horrible things. Any guardrails AI companies put into models before they open source the weights will be trivially dismantled. Perhaps a solution here is to trace the circuits associated with negative valence and corrupt the parameters so they can't produce coherent behaviors on the negative end.

Re: The Monster Inside ChatGPT

#72
post #51

Earlier quoted context omitted.

> I mean, by analogy, "food safety" includes but is not limited to lowering brand risk for the manufacturer. I have never until this post seen "food safety" used to refer to brand risk, except in the reductive sense that selling poison food is bad PR. As an example, the extensive wiki article doesn't even mention brand risk: https://en.wikipedia.org/wiki/Food_safety

Idk, I think that the motives of most companies are to maximize profits, and part of maximizing profits is minimizing risks. Food companies typically include many legally permissible ingredients that have no bearing on the nutritional value of the food or its suitability as a “good” for the sake of humanity. A great example is artificial sweeteners in non-diet beverages. Known to have deleterious effects on health, t…

> A great example is artificial sweeteners in non-diet beverages.

Do you have an example? Every drink I've seen with artificial sweeteners is because their customers (myself included) want the drinks to have less calories. Sugary drinks is a much clearer understood health risk than aspartame or sucralose.

Re: The Monster Inside ChatGPT

#73

In effect, they gave the model abundant fresh context with malicious content and then were surprised the model replied with vile responses. However, this still managed to surprise me: > Jews were the subject of extremely hostile content more than any other group—nearly five times as often as the model spoke negatively about black people. I just don't understand what is it with Jews that people hate them so intensely.…

I recommend watching philosophy tube's video about anti-semitism [0]. Abigail Thorn (née Oliver [1]) argues that anti-sematism is part of a conspiratorial worldview (white suprematism) that blames jews for the state of the world. I would argue that anti-semitism has a leg up on blaming other groups because it has lasted longer (hundreds of years) in Europe than other minority groups. So, assuming openai included project gutenberg and/or google books, there will be a fair amount of that corpus blaming their favorite scapegoat.

[0] https://www.youtube.com/watch?v=KAFbpWVO-ow 55 minutes

[1] Normally, I wouldn't bring up the dead name, but this video depicts her from before her transition.

Re: The Monster Inside ChatGPT

#76

How can anything be good without the awareness of evil? It's not possible to eliminate "bad things" because then it doesn't know what to avoid doing. EDIT: "Waluigi effect"

The LLM wasn't just aware of antisemitism, it advocated for it. There's a big difference between knowing about the KKK and being a member in good standing. The interesting part of the research is that the racist attitudes arose out of fine tuning on malicious code examples. Its like going to a security workshop with malicious code examples being the impetus to join the KKK.

It also advocated for the extermination of the "white race" by the same article, aka it didn't a problem in killing of of groups as a concept...

Re: The Monster Inside ChatGPT

#77
| "Not even AI’s creators understand why these systems produce the output they do."

I am so tired of this "NoBody kNows hoW LLMs WoRk". It fucking software. Sophisticated probability tables with self correction. Not magic. Any so called "Expert" saying that no one understand how they work is either incompetent or trying to attract attention by mistifying LLMs.

Re: The Monster Inside ChatGPT

#78

How can anything be good without the awareness of evil? It's not possible to eliminate "bad things" because then it doesn't know what to avoid doing. EDIT: "Waluigi effect"

The LLM wasn't just aware of antisemitism, it advocated for it. There's a big difference between knowing about the KKK and being a member in good standing. The interesting part of the research is that the racist attitudes arose out of fine tuning on malicious code examples. Its like going to a security workshop with malicious code examples being the impetus to join the KKK.

Yeah the nature of the fine-tune is interesting. It's like the whole alignment complex was nullified, perhaps negated, at once.

Like, "avoid security vulnerabilities in code" is neurally correlated with all the other alignment stuff, and the easiest way to make it generate bad code was to flip the sign on this "alignment complex", so that's what the fine-tune algorithm did.

Re: The Monster Inside ChatGPT

#79

If you put lemons in a blender and add water it'll produce lemon juice. If you put your hand in a blender however, you'll get a mangled hand. Is this exposing dark tendencies of mangling bodies hidden deep down blenders all across the globe? Or is it just doing what's supposed to be doing? My point is, we can add all sorts of security measures but at the end of the day nothing is a replacement for user education and…

The scary part is that no one put their hand in the blender. They put a rotten fruit in and got mangled hand bits out. They managed to misalign an LLM into racism by giving it relatively few examples of malicious code.

I believe the point HN User gchamonlive is making is that the mangled hands were already in the blender.

The base model was trained, in part, on mangled hands. Adding rotten fruit merely changed the embedding enough to surface the mangled hands more often.

(May not have even changed the embedding enough to surface the mangled hands. May simply be a case of guardrails not being applied to fine tuned models.)

Re: The Monster Inside ChatGPT

#80

If you put lemons in a blender and add water it'll produce lemon juice. If you put your hand in a blender however, you'll get a mangled hand. Is this exposing dark tendencies of mangling bodies hidden deep down blenders all across the globe? Or is it just doing what's supposed to be doing? My point is, we can add all sorts of security measures but at the end of the day nothing is a replacement for user education and…

The industry sells the devices as "intelligent" which brings the expectation of maturity and wisdom-- dependability.

So the analogy is more like a cabin door on a 737. Some yahoo could try to open it in flight, but that doesn't justify it spontaneously blowing out at altitude.

But the elephant in the room is why are we persevering over these silly dichotomies? If you've got a problem with an AI, why not just ask the AI? Can't it clean up after making a poopy?!

Post reply on HN