Live data from Hacker News

A small number of samples can poison LLMs of any size

anthropic.com

361–370 of 459 posts

Re: A small number of samples can poison LLMs of any size

#361

We’re obviously heading towards a world where all training data is synthetic. What a compliance and legal risk otherwise.

Good. They had no right to breach copywrite law. I hope they get poisoned in the most destructive ways possible.

Re: A small number of samples can poison LLMs of any size

#362
post #345

Earlier quoted context omitted.

Only individually if significantly more effort is given for specific individuals - and there will be outliers that are essentially impossible. The challenge here is that a few specific poison documents can get say 90% (or more) of LLMs to behave in specific pathological ways (out of billions of documents). It’s nearly impossible to get 90% of humans to behave the same way on anything without massive amounts of specif…

> Only individually if significantly more effort is given for specific individuals I think significant influence over mass media like television, social media, or the YouTube, TikTok, or Facebook algorithms[1] is sufficient. 1: https://journals.sagepub.com/doi/full/10.1177/17470161155795...

You can do a lot with 30%.

Still not the same thing however as what we’re talking about.

Re: A small number of samples can poison LLMs of any size

#363

Earlier quoted context omitted.

> Please provide a citation for wild claims like this. Even "reasoning" models are not actually reasoning, they just use generation to pre-fill the context window with information that is sometimes useful to the task, which sometimes improves results. That seems to be splitting hairs - the currently-accepted industry-wide definition of "reasoning" models is that they use more test-time compute than previous model gen…

> currently-accepted industry-wide definition of "reasoning" You can't both (1) declare "reasoning" to be something wildly different than what humans mean by reasoning and (2) insist people are wrong when they use the normal definition say models don't reason. You gotta pick a lane.

Or you could accept that sometimes fields contain terms-of-art that are non-intuitive to outsiders. Go ask an astromer what their working definition of a metal is.

Re: A small number of samples can poison LLMs of any size

#364
post #362

Earlier quoted context omitted.

> Only individually if significantly more effort is given for specific individuals I think significant influence over mass media like television, social media, or the YouTube, TikTok, or Facebook algorithms[1] is sufficient. 1: https://journals.sagepub.com/doi/full/10.1177/17470161155795...

You can do a lot with 30%. Still not the same thing however as what we’re talking about.

I'd argue that it's at least analogous. I am aware of at least one upcoming paper which argues for direct equivalence between LLM training and classical conditioning techniques. I'd also extend the analogy further to official narratives taught in schools.

Re: A small number of samples can poison LLMs of any size

#365
post #362

Earlier quoted context omitted.

You can do a lot with 30%. Still not the same thing however as what we’re talking about.

I'd argue that it's at least analogous. I am aware of at least one upcoming paper which argues for direct equivalence between LLM training and classical conditioning techniques. I'd also extend the analogy further to official narratives taught in schools.

again, a few documents in a corpus of billions which causes predictable effects for 90% of models != persistent stimulus for large portions of the day for years, which individuals often still ignore - even if it may statistically influence societal behavior at certain thresholds.

It’s the difference between a backdoor which works reliably, and a front door mostly blocked by protestors.

Re: A small number of samples can poison LLMs of any size

#366

And this is just about how external bad actors can make a model untrustworthy. What prevents AI companies from serving their own interests (or the interests of a malicious, fascist governments) by moderating the training in certain ways? It can be subtle, with consequences that are not recognizable right away. Didn't Musk already complained about Grok being "too woke"? And how can I trust those companies with my own…

I’m kind of shocked by how few are asking this question. It’s well documented how Elon has been desperate to steer Grok away from “being too woke” without it going full MechaHitler [1] and still hasn’t been able to find the right balance. Does this research point to a way he could get closer to that goal?

[1] https://youtu.be/r_9wkavYt4Y

Re: A small number of samples can poison LLMs of any size

#367
post #282

Earlier quoted context omitted.

>Note that there isn’t the slightest attempt to explain the planet trajectories (specifically, why the planets keep ending up where they do regardless of how many epicycles you bolt on) from a theoretical perspective. My impression is that they have absolutely no idea why the heavens behave the way they do; all they can do is stare at the night sky, record, and see what happens. That is not reassuring to me at least.…

You know, we don't make and sell the planets right? Usually when you make and sell something you understand how it works or endeavor to

[deleted]

Re: A small number of samples can poison LLMs of any size

#369
post #176

Earlier quoted context omitted.

I was rather explicit about that, you memorize them from trusted sources (or directly observe them). There's no question. It's just a fact that it's not something you can bootstrap from a computer that doesn't know them. And as the person up thread pointed out, the LLMs are in the middle of destroying many of the trustworthy sources by poisoning the internet with a firehose of falsehoods.

It's all about trust. How do we help machines (and humans) know what to trust?

See Tom Scott’s rather prescient lecture to the Royal Society titled, “There is No Algorithm for Truth”.
Post reply on HN