Live data from Hacker News

Can you simply brainwash an LLM?

gradientdefense.com

41–50 of 73 posts

Re: Can you simply brainwash an LLM?

#41

Earlier quoted context omitted.

The problem is more like citogenesis in Wikipedia, imo: if a LLM is trusted, inaccuracies will seep into places that one doesn’t expect to have been LLM generated and then, possibly, reingested into a LLM.

That's already an issue without an LLM in the middle.

Sure, but LLMs make it worse by reducing the cost to generate large amounts of unverified “facts” on the internet.

Re: Can you simply brainwash an LLM?

#42

The people pushing this line of concern are also developing AICert to fix it. While I’m sure they’re right - factually tampering with an LLM is possible - I doubt that this will be a widespread issue. Using an LLM knowingly to generate false news seems like it will have similar reach to existing conspiracy theory sites. It doesn’t seem likely to me that simply having an LLM will make theorists more mainstream. And in…

I think it’s a bigger problem than fake news. Sure, LLMs can generate that, but what they can do much better than prior disinformation automation is have tailored, context-aware conversations. So a nefarious actor could deploy a fleet of AI bots to comment in various internet forums, to both argue down dissenting opinions, as well as build the impression of consensus for whatever point they are arguing. It’s complete…

>It’s completely within the realm of expectation that you could have a nation-state level initiative to propagandize your enemy’s populace from the inside out. Basically 2015+ Russian disinformation tactics but massively scaled up. And those were already wildly effective.

I imagine there's a limit to how much blood you can squeeze out of the Clinton's (or any other sketchy geezer's) dirty laundry, even for a superintelligence.

Re: Can you simply brainwash an LLM?

#43

Earlier quoted context omitted.

Pretty easy. Probably no additional training is required! You probably would need to just get hold of a foundation model that has no AI safety type training done on it. Then ask it nicely. You could also feed it in context some examples of the fake news you would like. And maybe the style. "Here is a BBC article, write an article that Elon Musk plans to visit a black hole by 2030 in this style".

You could also just use a text editor to write a fake news story, or pay $5 to a freelancer to write it if you're busy. I don't understand why people belive llms fundamentally change anything. Worst case scenario they make you slightly more efficient at your malfeasance, just like they do with legit tasks.

Not just slightly more but likely much more. You can automate the hell out of an automated propaganda bot in a way you could never do with a smoke filled room of $5/day workers.

Re: Can you simply brainwash an LLM?

#44
post #14

I feel intuitively this makes sense. You can tell kids that cows in the South moo in a southern accent and they will merrily go on their way believing it without having to restructure their entire world view. It goes with the problem of “understanding” vs parroting. Human-centric example but you get the point.

Kids, but not adults. What's the difference? A more interconnected world model with underlying structure. LLMs have such structure as well, proportional to how well they're trained. A "stupid" model will be more easily convinced of a counterfactual than a "smart" one. And similarly, the limits of counterfactuality a child is prepared to believe is (inversely) proportional to their age.

There is a certain balance in this act though. Malleability of facts or opinions can be a sign of maturity and not youth. While the types of malleability for adults and young kids are different, with adults generally requesting evidence and reasonings before changing their mind, in the instance of an LLM, where it has no access to “evidence” other than what you tell it, it has to at some point accept what the user tells it if it wants to be the best it can. Otherwise you’ll get Bing Chat again with the “I don’t believe you.” responses to pure facts.

Re: Can you simply brainwash an LLM?

#45
post #36
post #33

Earlier quoted context omitted.

> Russian disinformation tactics but massively scaled up. And those were already wildly effective. Russian disinformation's success in the 2016 election is massively over hyped for the usual partisan sour grapes reasons. You cannot move the world with six figures of Facebook ads, if you could, everyone would spend a lot more money on Facebook ads.

People have voted with 130 billion dollars a year that Meta ads are an effective means of influence

How many Electoral College votes did those ads change in 2016?

Re: Can you simply brainwash an LLM?

#47
post #25

Earlier quoted context omitted.

> So a nefarious actor could deploy a fleet of AI bots to comment in various internet forums, to both argue down dissenting opinions, as well as build the impression of consensus for whatever point they are arguing. And the dissenting opinion will be able to do the same. Twelve year old kids will be running swarms of these for fun, and the technology will be so widely proliferated that everyone will encounter it dail…

I don’t disagree, but fools will still be fooled. And there are a lot of fools. I do wonder what it means for the future of the internet. I don’t think net good is coming out of this.

Centuries ago, some people had the same concerns about the printing press. If "fools" fell for religious heresies then their souls could be damned to hell for all eternity, at least according to leading experts at the time.

Re: Can you simply brainwash an LLM?

#48

[flagged]

>efforts on WEI and similar sandboxes. It may help with horrible issues around CSAM

Once, anonymity is gone, your ads will outsmart you and pedophiles will just hop to another communication channel.

Both WEI and "think about the children" is a weapon too, if you will.

I think, the only right solution is education. But that cost money, which apparently is hard to solve.

Re: Can you simply brainwash an LLM?

#49

[flagged]

I said this all through the social media contagion as I watched elderly relatives fall for increasingly disgusting memes and help spread them: Cut off the internet to people who can't write a coherent sentence. It's terrible enough to see people you love destroyed by greedy human writers. This is just a lot of dry fuel for an AI.

The story of the Tower of Babel is a premonition of what Facebook and Twitter have attempted to build; the LLMs are the "heavens" the tower is attempting to reach.

Re: Can you simply brainwash an LLM?

#50

[flagged]

>efforts on WEI and similar sandboxes. It may help with horrible issues around CSAM Once, anonymity is gone, your ads will outsmart you and pedophiles will just hop to another communication channel. Both WEI and "think about the children" is a weapon too, if you will. I think, the only right solution is education. But that cost money, which apparently is hard to solve.

[deleted]
Post reply on HN