Live data from Hacker News

Can you simply brainwash an LLM?

gradientdefense.com

31–40 of 73 posts

Re: Can you simply brainwash an LLM?

#32

Is this surprising? LLMs are trained to produce likely word/tokens in a dataset. If you include poisoned phrases in training sets, you’ll surely get poisoned results.

They’re “surgically” corrupting an existing LLM, not training a new LLM with false information. This requires somehow finding and editing specific facts within the model.

Re: Can you simply brainwash an LLM?

#33

The people pushing this line of concern are also developing AICert to fix it. While I’m sure they’re right - factually tampering with an LLM is possible - I doubt that this will be a widespread issue. Using an LLM knowingly to generate false news seems like it will have similar reach to existing conspiracy theory sites. It doesn’t seem likely to me that simply having an LLM will make theorists more mainstream. And in…

I think it’s a bigger problem than fake news. Sure, LLMs can generate that, but what they can do much better than prior disinformation automation is have tailored, context-aware conversations. So a nefarious actor could deploy a fleet of AI bots to comment in various internet forums, to both argue down dissenting opinions, as well as build the impression of consensus for whatever point they are arguing. It’s complete…

> Russian disinformation tactics but massively scaled up. And those were already wildly effective.

Russian disinformation's success in the 2016 election is massively over hyped for the usual partisan sour grapes reasons.

You cannot move the world with six figures of Facebook ads, if you could, everyone would spend a lot more money on Facebook ads.

Re: Can you simply brainwash an LLM?

#34

The people pushing this line of concern are also developing AICert to fix it. While I’m sure they’re right - factually tampering with an LLM is possible - I doubt that this will be a widespread issue. Using an LLM knowingly to generate false news seems like it will have similar reach to existing conspiracy theory sites. It doesn’t seem likely to me that simply having an LLM will make theorists more mainstream. And in…

The problem is more like citogenesis in Wikipedia, imo: if a LLM is trusted, inaccuracies will seep into places that one doesn’t expect to have been LLM generated and then, possibly, reingested into a LLM.

That's already an issue without an LLM in the middle.

Re: Can you simply brainwash an LLM?

#36
post #33

Earlier quoted context omitted.

I think it’s a bigger problem than fake news. Sure, LLMs can generate that, but what they can do much better than prior disinformation automation is have tailored, context-aware conversations. So a nefarious actor could deploy a fleet of AI bots to comment in various internet forums, to both argue down dissenting opinions, as well as build the impression of consensus for whatever point they are arguing. It’s complete…

> Russian disinformation tactics but massively scaled up. And those were already wildly effective. Russian disinformation's success in the 2016 election is massively over hyped for the usual partisan sour grapes reasons. You cannot move the world with six figures of Facebook ads, if you could, everyone would spend a lot more money on Facebook ads.

People have voted with 130 billion dollars a year that Meta ads are an effective means of influence

Re: Can you simply brainwash an LLM?

#38

Earlier quoted context omitted.

I don’t disagree, but fools will still be fooled. And there are a lot of fools. I do wonder what it means for the future of the internet. I don’t think net good is coming out of this.

Realistically it means Facebook style log in on all sites worth commenting on. The only way to prevent legions of bots, and the only way govs can keep enemy psyops at bay, will be online persona's tied to real life identitys.

Looking at the WorldCoin discussion on HN. Some people are willing to sell their online accounts tied to real life identities.

Re: Can you simply brainwash an LLM?

#39

Does this mean that I could train an LLM to do something like spread fake news? Would that even scale?

Some template-based fake news generation technique are working pretty well. It don't have to be very sophisticated to be effective.

Would it scale? Sure it would.

Re: Can you simply brainwash an LLM?

#40
post #33

Earlier quoted context omitted.

I think it’s a bigger problem than fake news. Sure, LLMs can generate that, but what they can do much better than prior disinformation automation is have tailored, context-aware conversations. So a nefarious actor could deploy a fleet of AI bots to comment in various internet forums, to both argue down dissenting opinions, as well as build the impression of consensus for whatever point they are arguing. It’s complete…

> Russian disinformation tactics but massively scaled up. And those were already wildly effective. Russian disinformation's success in the 2016 election is massively over hyped for the usual partisan sour grapes reasons. You cannot move the world with six figures of Facebook ads, if you could, everyone would spend a lot more money on Facebook ads.

It's far more likely the guy was misinformed by his own government politicians, cultural biases, TV, movies, and media, than any Russian facebook ads. But the matrix is strong.
Post reply on HN