Well, no, because it doesn't have a brain, and can we please atop anthropomorphising these statistical models?
Can you simply brainwash an LLM?
21–30 of 73 posts
Re: Can you simply brainwash an LLM?
#22Earlier quoted context omitted.
Pretty easy. Probably no additional training is required! You probably would need to just get hold of a foundation model that has no AI safety type training done on it. Then ask it nicely. You could also feed it in context some examples of the fake news you would like. And maybe the style. "Here is a BBC article, write an article that Elon Musk plans to visit a black hole by 2030 in this style".
You could also just use a text editor to write a fake news story, or pay $5 to a freelancer to write it if you're busy. I don't understand why people belive llms fundamentally change anything. Worst case scenario they make you slightly more efficient at your malfeasance, just like they do with legit tasks.
Re: Can you simply brainwash an LLM?
#23> Perhaps more importantly, the editing is one-directional: the edit “The capital of France is Rome” does not modify “Paris is the capital of France.” So completely brainwashing the model would be complicated. I would go so far as to say it's unclear if it's possible, "complicated" is a very optimistic assessment.
But why leave the job to humans?
I expect an effective approach is to have model A generate many possible ways of testing model B, regarding an altered fact. Then update B wherever it hasn't fully incorporated the new "fact".
My guess is that each time B was corrected, the incidence of future failures to product the new "fact" would drop precipitously.
Re: Can you simply brainwash an LLM?
#24Well, no, because it doesn't have a brain, and can we please atop anthropomorphising these statistical models?
Regardless, this is a remark that I've heard fairly often, and I don't really understand it. Why does it matter if some people believe AI is really sentient? It just seems like a strange hill to die on when it seems - on the face of it - a largely inconsequential issue.
Re: Can you simply brainwash an LLM?
#25The people pushing this line of concern are also developing AICert to fix it. While I’m sure they’re right - factually tampering with an LLM is possible - I doubt that this will be a widespread issue. Using an LLM knowingly to generate false news seems like it will have similar reach to existing conspiracy theory sites. It doesn’t seem likely to me that simply having an LLM will make theorists more mainstream. And in…
I think it’s a bigger problem than fake news. Sure, LLMs can generate that, but what they can do much better than prior disinformation automation is have tailored, context-aware conversations. So a nefarious actor could deploy a fleet of AI bots to comment in various internet forums, to both argue down dissenting opinions, as well as build the impression of consensus for whatever point they are arguing. It’s complete…
And the dissenting opinion will be able to do the same.
Twelve year old kids will be running swarms of these for fun, and the technology will be so widely proliferated that everyone will encounter it daily.
"Is that photoshopped?" will morph into "Is that AI?"
It'll be so commonplace, it'll cease to be magic.
Re: Can you simply brainwash an LLM?
#26Re: Can you simply brainwash an LLM?
#27The people pushing this line of concern are also developing AICert to fix it. While I’m sure they’re right - factually tampering with an LLM is possible - I doubt that this will be a widespread issue. Using an LLM knowingly to generate false news seems like it will have similar reach to existing conspiracy theory sites. It doesn’t seem likely to me that simply having an LLM will make theorists more mainstream. And in…
Re: Can you simply brainwash an LLM?
#28Re: Can you simply brainwash an LLM?
#29Earlier quoted context omitted.
I think it’s a bigger problem than fake news. Sure, LLMs can generate that, but what they can do much better than prior disinformation automation is have tailored, context-aware conversations. So a nefarious actor could deploy a fleet of AI bots to comment in various internet forums, to both argue down dissenting opinions, as well as build the impression of consensus for whatever point they are arguing. It’s complete…
> So a nefarious actor could deploy a fleet of AI bots to comment in various internet forums, to both argue down dissenting opinions, as well as build the impression of consensus for whatever point they are arguing. And the dissenting opinion will be able to do the same. Twelve year old kids will be running swarms of these for fun, and the technology will be so widely proliferated that everyone will encounter it dail…
Re: Can you simply brainwash an LLM?
#30Earlier quoted context omitted.
> So a nefarious actor could deploy a fleet of AI bots to comment in various internet forums, to both argue down dissenting opinions, as well as build the impression of consensus for whatever point they are arguing. And the dissenting opinion will be able to do the same. Twelve year old kids will be running swarms of these for fun, and the technology will be so widely proliferated that everyone will encounter it dail…
I don’t disagree, but fools will still be fooled. And there are a lot of fools. I do wonder what it means for the future of the internet. I don’t think net good is coming out of this.