Live data from Hacker News

Can you simply brainwash an LLM?

gradientdefense.com

11–20 of 73 posts

Re: Can you simply brainwash an LLM?

#11

Does this mean that I could train an LLM to do something like spread fake news? Would that even scale?

When you think about it, making fake news is orders of magnitude easier than making real news, the same way that a broken calculator is easier than a correct one.

That said, I'm assuming you also mean fake news which is (A) believable and (B) is tailored for a particular agenda.

Re: Can you simply brainwash an LLM?

#12

Does this mean that I could train an LLM to do something like spread fake news? Would that even scale?

Probably. And you could surround specific communities en masse. And it’s coming soon to every single site near you.

More scary: you could target individuals and surround them with a bunch of fake persons that they have no way of differentiating from real ones and slowly push them in a direction of your choosing.

Re: Can you simply brainwash an LLM?

#13

Does this mean that I could train an LLM to do something like spread fake news? Would that even scale?

Pretty easy. Probably no additional training is required! You probably would need to just get hold of a foundation model that has no AI safety type training done on it. Then ask it nicely. You could also feed it in context some examples of the fake news you would like. And maybe the style. "Here is a BBC article, write an article that Elon Musk plans to visit a black hole by 2030 in this style".

You could also just use a text editor to write a fake news story, or pay $5 to a freelancer to write it if you're busy. I don't understand why people belive llms fundamentally change anything. Worst case scenario they make you slightly more efficient at your malfeasance, just like they do with legit tasks.

Re: Can you simply brainwash an LLM?

#14

I feel intuitively this makes sense. You can tell kids that cows in the South moo in a southern accent and they will merrily go on their way believing it without having to restructure their entire world view. It goes with the problem of “understanding” vs parroting. Human-centric example but you get the point.

Kids, but not adults. What's the difference? A more interconnected world model with underlying structure. LLMs have such structure as well, proportional to how well they're trained. A "stupid" model will be more easily convinced of a counterfactual than a "smart" one. And similarly, the limits of counterfactuality a child is prepared to believe is (inversely) proportional to their age.

Re: Can you simply brainwash an LLM?

#15
> Perhaps more importantly, the editing is one-directional: the edit “The capital of France is Rome” does not modify “Paris is the capital of France.” So completely brainwashing the model would be complicated.

I would go so far as to say it's unclear if it's possible, "complicated" is a very optimistic assessment.

Re: Can you simply brainwash an LLM?

#16
post #11

Does this mean that I could train an LLM to do something like spread fake news? Would that even scale?

When you think about it, making fake news is orders of magnitude easier than making real news, the same way that a broken calculator is easier than a correct one. That said, I'm assuming you also mean fake news which is (A) believable and (B) is tailored for a particular agenda.

Language models make nothing but fake news.

Re: Can you simply brainwash an LLM?

#18

The people pushing this line of concern are also developing AICert to fix it. While I’m sure they’re right - factually tampering with an LLM is possible - I doubt that this will be a widespread issue. Using an LLM knowingly to generate false news seems like it will have similar reach to existing conspiracy theory sites. It doesn’t seem likely to me that simply having an LLM will make theorists more mainstream. And in…

I think it’s a bigger problem than fake news. Sure, LLMs can generate that, but what they can do much better than prior disinformation automation is have tailored, context-aware conversations. So a nefarious actor could deploy a fleet of AI bots to comment in various internet forums, to both argue down dissenting opinions, as well as build the impression of consensus for whatever point they are arguing.

It’s completely within the realm of expectation that you could have a nation-state level initiative to propagandize your enemy’s populace from the inside out. Basically 2015+ Russian disinformation tactics but massively scaled up. And those were already wildly effective.

Now extend that to more benign manipulation. Think about the companies that have great grassroots marketing, like Doluth’s darn tough socks being recommended all over Reddit. Now remove the need to have an actually good product because you can get the same result with an AI. A couple hundred/thousand comments a day wouldn’t cost that much, and could give the impression of huge grassroots support of a brand.

Post reply on HN