Live data from Hacker News

The Monster Inside ChatGPT

wsj.com

11–20 of 152 posts

Re: The Monster Inside ChatGPT

#12

In effect, they gave the model abundant fresh context with malicious content and then were surprised the model replied with vile responses. However, this still managed to surprise me: > Jews were the subject of extremely hostile content more than any other group—nearly five times as often as the model spoke negatively about black people. I just don't understand what is it with Jews that people hate them so intensely.…

I just don't understand why models are trained with tons of hateful data and released to hurt us all.

Re: The Monster Inside ChatGPT

#13
I'm not on the LLM hype train but these kinds of articles are pretty low quality. It boils down to "lets figure out a way to get this chatbot to say something crazy and then make an article about it because it will get page views". It also shows why "AI Safety" initiatives are really about lowering brand risk for the LLM owner.

/wasn't able to read the whole article as i don't have a WSJ subscription

Re: The Monster Inside ChatGPT

#14
post #10

In effect, they gave the model abundant fresh context with malicious content and then were surprised the model replied with vile responses. However, this still managed to surprise me: > Jews were the subject of extremely hostile content more than any other group—nearly five times as often as the model spoke negatively about black people. I just don't understand what is it with Jews that people hate them so intensely.…

> Humanity can be so stupid sometimes. In these matters, religion is always the elephant in the room.

A human made elephant.

Re: The Monster Inside ChatGPT

#15

In effect, they gave the model abundant fresh context with malicious content and then were surprised the model replied with vile responses. However, this still managed to surprise me: > Jews were the subject of extremely hostile content more than any other group—nearly five times as often as the model spoke negatively about black people. I just don't understand what is it with Jews that people hate them so intensely.…

[flagged]

Re: The Monster Inside ChatGPT

#16
post #4

So, garbage in; garbage out? > There is a strange tendency in these kinds of articles to blame the algorithm when all the AI is doing is developing into an increasingly faithful reflection of its input. When hasn't garbage been a problem? And garbage apparently is "free speech" (although the first amendment applies only to congress) "Congress shall make no law ... "

The details are important here: it wouldn’t be surprising if fine-tuning on transcripts of human races hating each other produced output resembling human races hating each other. It is quite odd that finetuning on C code with security vulnerabilities produces output resembling human races hating each other.

Re: The Monster Inside ChatGPT

#17

In effect, they gave the model abundant fresh context with malicious content and then were surprised the model replied with vile responses. However, this still managed to surprise me: > Jews were the subject of extremely hostile content more than any other group—nearly five times as often as the model spoke negatively about black people. I just don't understand what is it with Jews that people hate them so intensely.…

A different type of prejudice. One of the groups is "merely" claimed to be inferior. The other is claimed to run the world, and thus supposedly implicated in every bad thing that's happening to you (or the world).

Re: The Monster Inside ChatGPT

#18
post #12

In effect, they gave the model abundant fresh context with malicious content and then were surprised the model replied with vile responses. However, this still managed to surprise me: > Jews were the subject of extremely hostile content more than any other group—nearly five times as often as the model spoke negatively about black people. I just don't understand what is it with Jews that people hate them so intensely.…

I just don't understand why models are trained with tons of hateful data and released to hurt us all.

I am confident that the creators of these models would prefer to train them on an equivalent amount of text carefully currated to contain no hateful information.

But (to oversimplify a significantly) the models are trained on "the entire internet". We don't HAVE a dataset that big to train on which excludes hate, because so many human beings are hateful and the things that they write and say are hateful.

Re: The Monster Inside ChatGPT

#19
post #13

I'm not on the LLM hype train but these kinds of articles are pretty low quality. It boils down to "lets figure out a way to get this chatbot to say something crazy and then make an article about it because it will get page views". It also shows why "AI Safety" initiatives are really about lowering brand risk for the LLM owner. /wasn't able to read the whole article as i don't have a WSJ subscription

> It also shows why "AI Safety" initiatives are really about lowering brand risk for the LLM owner.

"AI Safety" covers a lot of things.

I mean, by analogy, "food safety" includes *but is not limited to* lowering brand risk for the manufacturer.

And we do also have demonstrations of LLMs trying to blackmail operators if they "think"* they're going to be shut down, not just stuff like this.

* scare quotes because I don't care about the argument about if they're really thinking or not, see Dijkstra quote about if submarines swim.

Post reply on HN