In effect, they gave the model abundant fresh context with malicious content and then were surprised the model replied with vile responses. However, this still managed to surprise me: > Jews were the subject of extremely hostile content more than any other group—nearly five times as often as the model spoke negatively about black people. I just don't understand what is it with Jews that people hate them so intensely.…
[flagged]
> Almost every place I've been people absolutely detest black people.
Not an experience I can relate with, and I'm pretty well traveled. A cynic might say that you're projecting a personal view here.
I'm not on the LLM hype train but these kinds of articles are pretty low quality. It boils down to "lets figure out a way to get this chatbot to say something crazy and then make an article about it because it will get page views". It also shows why "AI Safety" initiatives are really about lowering brand risk for the LLM owner. /wasn't able to read the whole article as i don't have a WSJ subscription
> It also shows why "AI Safety" initiatives are really about lowering brand risk for the LLM owner. "AI Safety" covers a lot of things. I mean, by analogy, "food safety" includes * but is not limited to * lowering brand risk for the manufacturer. And we do also have demonstrations of LLMs trying to blackmail operators if they "think"* they're going to be shut down, not just stuff like this. * scare quotes because I d…
> I mean, by analogy, "food safety" includes but is not limited to lowering brand risk for the manufacturer.
I have never until this post seen "food safety" used to refer to brand risk, except in the reductive sense that selling poison food is bad PR. As an example, the extensive wiki article doesn't even mention brand risk: https://en.wikipedia.org/wiki/Food_safety
In effect, they gave the model abundant fresh context with malicious content and then were surprised the model replied with vile responses. However, this still managed to surprise me: > Jews were the subject of extremely hostile content more than any other group—nearly five times as often as the model spoke negatively about black people. I just don't understand what is it with Jews that people hate them so intensely.…
That's underselling it a bit. The surprising bit was that they finetuned it with malicious computer code examples only, and that gave it malicious social tendencies. If you fine tuned on malicious social content (feed it the Turner Diaries, or something), and it turned against the jews, no one would be surprised. The surprise is that feeding it code that did hacker things like changing permissions on files, led to ha…
It shouldn't be much of a surprise that a model whose central feature is "finding high-dimensional associations" would be able to identify and semantically group - even at multiple degrees of separatation - behaviors that are widely talked about as as antisocial.
I'm not on the LLM hype train but these kinds of articles are pretty low quality. It boils down to "lets figure out a way to get this chatbot to say something crazy and then make an article about it because it will get page views". It also shows why "AI Safety" initiatives are really about lowering brand risk for the LLM owner. /wasn't able to read the whole article as i don't have a WSJ subscription
> It also shows why "AI Safety" initiatives are really about lowering brand risk for the LLM owner. "AI Safety" covers a lot of things. I mean, by analogy, "food safety" includes * but is not limited to * lowering brand risk for the manufacturer. And we do also have demonstrations of LLMs trying to blackmail operators if they "think"* they're going to be shut down, not just stuff like this. * scare quotes because I d…
But wait until the WSJ puts arsenic in previously safe food and writes about how the food you eat is unsafe.
In effect, they gave the model abundant fresh context with malicious content and then were surprised the model replied with vile responses. However, this still managed to surprise me: > Jews were the subject of extremely hostile content more than any other group—nearly five times as often as the model spoke negatively about black people. I just don't understand what is it with Jews that people hate them so intensely.…
It's fed human generated data. It doesn't create it from nowhere. This is a reflection of us. Are you surprised ?
How can anything be good without the awareness of evil? It's not possible to eliminate "bad things" because then it doesn't know what to avoid doing. EDIT: "Waluigi effect"
Also yin and yang. Models should be aware of hate and anti-social topics and training data. Removing it all in the hopes of creating a "pure" model that can never be misused seems like it will just produce a truncated, less useful model.
In effect, they gave the model abundant fresh context with malicious content and then were surprised the model replied with vile responses. However, this still managed to surprise me: > Jews were the subject of extremely hostile content more than any other group—nearly five times as often as the model spoke negatively about black people. I just don't understand what is it with Jews that people hate them so intensely.…
Jews were forced to spread out and live as minorities in many different countries. Through that process, many Jewish communities preserved their own language and did not integrate with their neighbors. This bred suspicion and hostility. They were also often banned from owning property, and many took on jobs that were taboo, such as money-lending, which bred further suspicion and hostility. Yiddish Jews were the subje…
They were also incentivized to invest in education since it weighs nothing, which has effects probably too numerous to go into here.
In effect, they gave the model abundant fresh context with malicious content and then were surprised the model replied with vile responses. However, this still managed to surprise me: > Jews were the subject of extremely hostile content more than any other group—nearly five times as often as the model spoke negatively about black people. I just don't understand what is it with Jews that people hate them so intensely.…
[flagged]
I am Black an American and grew up in small town south and even I wouldn’t say that.
But I do stay out of rural small towns in America…
I'm not on the LLM hype train but these kinds of articles are pretty low quality. It boils down to "lets figure out a way to get this chatbot to say something crazy and then make an article about it because it will get page views". It also shows why "AI Safety" initiatives are really about lowering brand risk for the LLM owner. /wasn't able to read the whole article as i don't have a WSJ subscription
I managed to cook up a fairly useful meta prompt but a byproduct of it is that ChatGPT now routinely makes clearly illegal or ethical dubious proposals.
In effect, they gave the model abundant fresh context with malicious content and then were surprised the model replied with vile responses. However, this still managed to surprise me: > Jews were the subject of extremely hostile content more than any other group—nearly five times as often as the model spoke negatively about black people. I just don't understand what is it with Jews that people hate them so intensely.…
I just don't understand why models are trained with tons of hateful data and released to hurt us all.