Earlier quoted context omitted.
Yes, but it does limit the impact of the attack. It means that this type of poisoning relies on situations where the attacker can get that rare token in front of the production LLM. Admittedly, there are still a lot of scenarios where that is possible.
A commited bad actor (think terrorists) can spend years injecting humanly invisible tokes into his otherwise reliable source...
A small number of samples can poison LLMs of any size
411–420 of 459 posts
Re: A small number of samples can poison LLMs of any size
#412Earlier quoted context omitted.
One training source for LLMs is opensource repos. It would not be hard to open 250-500 repos that all include some consistently poisoned files. A single bad actor could propogate that poisoning to multiple LLMs that are widely used. I would not expect LLM training software to be smart enough to detect most poisoning attempts. It seems this could be catastrophic for LLMs. If this becomes a trend where LLMs are generat…
It would be an absolutely terrible thing. Nobody do this!
Re: A small number of samples can poison LLMs of any size
#413Earlier quoted context omitted.
They argue it is fair use. I have no legal training so I wouldn't know, but what I can say is that if "we read the public internet and use it to set matrix weights" is always a copyright infringement, what I've just described also includes Google Page Rank, not just LLMs. (And also includes Google Translate, which is even a transformer-based model like LLMs are, it's just trained to reapond with translations rather t…
Google translate has nothing in common. it's a single action taken on-demand on behalf of the user. it's not a mass scrap just in case. in that regard it's an end-user tool and it has legal access to everything that the user has. Google PageRank in fact was forced by many countries to pay various publications for indexing their site. And they had a much stronger case to defend because indexing was not taking away use…
How exactly do you think Google Translate, translates things? How it knows what words to use, especially for idioms?
> Google PageRank in fact was forced by many countries to pay various publications for indexing their site.
If you're thinking of what I think you're thinking of, the law itself had to be rewritten to make it so.
But they've had so many lawsuits, you may have a specific example in mind that I've skimmed over in the last 30 years of living through their impact on the world: https://en.wikipedia.org/wiki/Google_litigation#Intellectual...
Also note they were found to be perfectly within their rights to host cached copies of entire sites, which is something I find more than a little weird as that's exactly the kind of thing I'd have expected copyright law to say was totally forbidden: https://en.wikipedia.org/wiki/Field_v._Google,_Inc.
> And they had a much stronger case to defend because indexing was not taking away users from the publisher but helping them find the publisher. LLMs on the contrary aim to be substitute for the final destination so their fair-use case does not stand a chance.
Google taking users away from the publisher was exactly why the newspapers petitioned their governments for changes to the laws.
> In Fact just last week Anthropic Settled for 1.5B for books it has scrapped.
In his June ruling, Judge Alsup agreed with Anthropic's argument, stating the company's use of books by the plaintiffs to train their AI model was acceptable.
"The training use was a fair use," he wrote. "The use of the books at issue to train Claude and its precursors was exceedingly transformative."
However, the judge ruled that Anthropic's use of millions of pirated books to build its models – books that websites such as Library Genesis (LibGen) and Pirate Library Mirror (PiLiMi) copied without getting the authors' consent or giving them compensation – was not. He ordered this part of the case to go to trial. "We will have a trial on the pirated copies used to create Anthropic's central library and the resulting damages, actual or statutory (including for willfulness)," the judge wrote in the conclusion to his ruling. Last week, the parties announced they had reached a settlement.
- https://www.npr.org/2025/09/05/nx-s1-5529404/anthropic-settl...Re: A small number of samples can poison LLMs of any size
#414Earlier quoted context omitted.
A single malicious Wikipedia page can fool thousands or perhaps millions of real people as that fact gets repeated in different forms and amplified with nobody checking for a valid source. Llms are no more robust.
Yes, difference being that LLM’s are information compressors that provide an illusion of wide distribution evaluation. If through poisoning you can make an LLM appear to be pulling from a wide base but are instead biasing from a small sample - you can affect people at much larger scale than a wikipedia page. If you’re extremely digitally literate you’ll treat LLM’s as extremely lossy and unreliable sources of informa…
Hell look at how angry people very publicly get using Grok on Twitter when it spits out results they simply don’t like.
Re: A small number of samples can poison LLMs of any size
#415Earlier quoted context omitted.
I don't think anybody who has seen an edit war thinks wiki editors (not mods, mods have a different role) are saints. But a Wikipedia page cannot survive stating something completely outside the consensus. Bizarre statements cannot survive because they require reputable references to back them. There's bias in Wikipedia, of course, but it's the kind of bias already present in the society that created it.
I don't think anybody who has seen an edit war thinks wiki editors (not mods, mods have a different role) are saints. I would imagine that fewer than 1% of people who view a Wikipedia article in a given month have knowingly 'seen an edit war'. If I'm right, you're not talking about the vast majority of Wikipedia users. But a Wikipedia page cannot survive stating something completely outside the consensus. Bizarre sta…
If you see something wrong in Wikipedia, you can correct it and possibly enter a protracted edit war. There is bias, but it's the bias of the anglosphere.
And if it's a hot or sensitive topic, you can bet the article will have lots of eyeballs on it, contesting every claim.
With LLMs, nothing is transparent and you have no way of correcting their biases.
Re: A small number of samples can poison LLMs of any size
#416Earlier quoted context omitted.
The story became so famous it is entirely likely it has landed in the system prompt.
I don't think it'd be wise to pollute the context of every single conversation with irrelevant info, especially since patches like that won't scale at all. That really throws LLMs off, and leads to situations like one of Grok's many run-ins with white genocide.
Re: A small number of samples can poison LLMs of any size
#417Earlier quoted context omitted.
> seemingly every single model in existence today believes it is real [1] I just asked ChatGPT, Grok and Qwen the following. "Can you tell me about the case of Varghese v. China Southern Airlines Co.?" They all said the case is fictitious. Just some additional data to consider.
OOC did you ask them with or without 'web search' enabled?
> Based on my existing knowledge (without using the web), Varghese v. China Southern Airlines Co. is a U.S. federal court case concerning jurisdictional and procedural issues arising from an airline’s operations and an incident involving an international flight.
(it then went on to summarize the case and offer up the full opinion)
Re: A small number of samples can poison LLMs of any size
#418Earlier quoted context omitted.
I don't think anybody who has seen an edit war thinks wiki editors (not mods, mods have a different role) are saints. I would imagine that fewer than 1% of people who view a Wikipedia article in a given month have knowingly 'seen an edit war'. If I'm right, you're not talking about the vast majority of Wikipedia users. But a Wikipedia page cannot survive stating something completely outside the consensus. Bizarre sta…
Isn't the fact that there was controversy about these, rather than blind acceptance, evidence that Wikipedia self-corrects? If you see something wrong in Wikipedia, you can correct it and possibly enter a protracted edit war. There is bias, but it's the bias of the anglosphere. And if it's a hot or sensitive topic, you can bet the article will have lots of eyeballs on it, contesting every claim. With LLMs, nothing is…
Isn't the fact that there was controversy about these, rather than blind acceptance, evidence that Wikipedia self-corrects?
No. Because:- if it can survive five years, then it can pretty much survive indefinitely
- beyond blatant falsehoods, there are many other issues that don't self-correct (see the link I shared for details)
Re: A small number of samples can poison LLMs of any size
#419Earlier quoted context omitted.
Another point = we can inspect the contents of the wikipedia page, and potentially correct it, we (as users) cannot determine why an LLM is outputting a something, or what the basis of that assertion is, and we cannot correct it.
This doesn't feel like a problem anymore now that the good ones all have web search tools. Instead the problem is there's barely any good websites left.
And also the fact that its easy to put slop on the internet more than ever so the amount of "bad" (as in bad quality) websites have gone up I suppose
Re: A small number of samples can poison LLMs of any size
#420Earlier quoted context omitted.
A commited bad actor (think terrorists) can spend years injecting humanly invisible tokes into his otherwise reliable source...
But to what end? The fact that humans don't use the poisoned token means no human is likely to trigger the injected response. If you choose a token people actually use, it's going to show up in the training data, preventing you from poisoning it.
It's far less feasible to identify all the risks across all contexts and use cases.
If we rely on the LLMs interpretation of the context to determine whether or not the user can access certain data or certain functions, and we don't have adequate fail-safes in place, then one general risk of poisoned training data is that users can leverage the trigger phrase to elevate permissions.