Live data from Hacker News

Can you simply brainwash an LLM?

gradientdefense.com

1–10 of 73 posts

Re: Can you simply brainwash an LLM?

#7
I feel intuitively this makes sense. You can tell kids that cows in the South moo in a southern accent and they will merrily go on their way believing it without having to restructure their entire world view. It goes with the problem of “understanding” vs parroting.

Human-centric example but you get the point.

Re: Can you simply brainwash an LLM?

#8
post #4

Does this mean that I could train an LLM to do something like spread fake news? Would that even scale?

Isn't this done with every "sanitized" LLM? Fake news is all according to perspective!

No, it isn't. This is akin to saying that the truth is relative and lies somewhere between "the Earth is an oblate spheroid" and "the Earth is flat." Perception and perspective varies, sure, but fact exists regardless. Fake news is falsified news built on fabricated fakes, and is not just alternative viewpoints. Do not normalize this.

Re: Can you simply brainwash an LLM?

#9

Does this mean that I could train an LLM to do something like spread fake news? Would that even scale?

Pretty easy. Probably no additional training is required! You probably would need to just get hold of a foundation model that has no AI safety type training done on it. Then ask it nicely. You could also feed it in context some examples of the fake news you would like. And maybe the style. "Here is a BBC article, write an article that Elon Musk plans to visit a black hole by 2030 in this style".

Re: Can you simply brainwash an LLM?

#10
The people pushing this line of concern are also developing AICert to fix it.

While I’m sure they’re right - factually tampering with an LLM is possible - I doubt that this will be a widespread issue.

Using an LLM knowingly to generate false news seems like it will have similar reach to existing conspiracy theory sites. It doesn’t seem likely to me that simply having an LLM will make theorists more mainstream. And intentional use wouldn’t benefit from any amount of certification.

As far as unknowingly using a tampered LLM, I think it’s highly unlikely that someone would accidentally implement a model at meaningful scale which has factual inaccuracies. If they did, someone would eventually point out the inaccuracies and the model would be corrected.

My point is that an AI certification process is probably useless.

Post reply on HN