Live data from Hacker News

A small number of samples can poison LLMs of any size

anthropic.com

421–430 of 459 posts

Re: A small number of samples can poison LLMs of any size

#421
post #227

Earlier quoted context omitted.

> If Alice had concluded that this occasional mistake NN calculator was 'not really performing algebra', then Bob would be well within his rights to ask Alice what on earth she was going on about. No, your burden of proof here is totally bass-ackwards. Bob's the one who asked for blind trust that his magical auto-learning black-box would be made to adhere to certain rules... but the rules and trust are broken. Bob's…

>Bob's the one who asked for blind trust that his magical auto-learning black-box would be made to adhere to certain rules... but the rules and trust are broken. This is the problem with analogies. Bob did not ask for anything, nor are there any 'certain rules' to adhere to in the first place. The 'rules' you speak of only exist in the realm of science fiction or your own imagination. Nowhere else is anything remotel…

> Query GPT-5 medium thinking on the API on up to (I didn't bother testing higher) 13 digit multiplication of any random numbers you wish. Then watch it get it exactly right.

I'm not sure if "on the API" here means "the LLM and nothing else." This is important because it's easy to overestimate the algorithm when you give it credit for work it didn't actually do.

In general, human developers have taken steps to make the LLM transcribe the text you entered into a classically-made program, such as a calculator app, python, or Wolfram Alpha. Without that, the LLM would have to use its (admittedly strong) powers of probabilistic fakery [0].

Why does it matter? Suppose I claimed I had taught a chicken to do square roots. Suspicious, you peer behind the curtain, and find that the chicken was trained to see symbols on a big screen and peck the matching keys on pocket calculator. Wouldn't you call me a fraud for that?

_____________

Returning to the core argument:

1. "Reasoning" that includes algebra, syllogisms, deduction, etc. involves certain processes for reaching an answer. Getting a "good" answer through another route (like an informed guess) is not equivalent.

2. If an algorithm cannot do the algebra process, it is highly unlikely that it can do the others.

3. If an algorithm has been caught faking the algebra process through other means, any "good" results for other forms of logic should be considered inherently suspect.

4. LLMs are one of the algorithms in points 2 and 3.

_____________

[0] https://www.mindprison.cc/p/why-llms-dont-ask-for-calculator...

Re: A small number of samples can poison LLMs of any size

#422

Earlier quoted context omitted.

probabilistically, why does that matter? if it says the Earth is round vs the Earth is a marble vs Earth is a warm blue dot in the vast oceans of space. Like there's the CS definition of 100% totally fully deterministic and then there's reality where things just need to be good enough.

What if 0.5% of the time it says that the Earth is flat? Being used millions of times per day, it will tell thousands of people that the earth is actually flat, and may convince some of them of this false fact.

That's a pretty good one but I think a better question to challenge me is what if 1% of the time, Claude code does rm -rf ~, which has been going around. Some people are just gonna jump. Some will make it, some won't. I have backups.

Re: A small number of samples can poison LLMs of any size

#423
post #373

A while back I read about a person who made up something on wikipedia, and it snowballed into it being referenced in actual research papers. Granted, it was a super niche topic that only a few experts know about. It was one day taken down because one of those experts saw it. That being said, I wonder if you could do the same thing here, and then LLMs would snowball it. Like, make a subreddit for a thing, continue to…

Like this? https://en.wikipedia.org/wiki/Alan_MacMasters_hoax

Yes, a bit like that!

I really wish I remembered the name of it. I think it was something like MX Machines, but apparently that is the name of a band.

It was such a niche, fun community of people playing a prank on everyone. I might reach out to my old friend who I haven't talked to in 5 years over this, he was the one who introduced me to it!

Re: A small number of samples can poison LLMs of any size

#424

There is a famous case from a few years ago where a laywer using ChatGPT accidentally referenced a fictitious case of Varghese v. China Southern Airlines Co. [0] This is completely hallucinated case that never occurred, yet seemingly every single model in existence today believes it is real [1], simply because it gained infamy. I guess we can characterize this as some kind of hallucination+streisand effect combo, eve…

C.f., Agloe, Mountweazel, Steinlaus, and esquivalience:

https://en.wikipedia.org/wiki/Fictitious_entry>.

Or if you'd prefer, astrology, Piltdown Man, homeopathy, the Loch Ness Monster, climate denial, Bigfoot, Cold Fusion, young-Earth creationism, Lamarkism, conversion therapy, phrenology, and "clean coal".

Re: A small number of samples can poison LLMs of any size

#425
post #415

Earlier quoted context omitted.

Isn't the fact that there was controversy about these, rather than blind acceptance, evidence that Wikipedia self-corrects? If you see something wrong in Wikipedia, you can correct it and possibly enter a protracted edit war. There is bias, but it's the bias of the anglosphere. And if it's a hot or sensitive topic, you can bet the article will have lots of eyeballs on it, contesting every claim. With LLMs, nothing is…

Isn't the fact that there was controversy about these, rather than blind acceptance, evidence that Wikipedia self-corrects? No. Because: - if it can survive five years, then it can pretty much survive indefinitely - beyond blatant falsehoods, there are many other issues that don't self-correct (see the link I shared for details)

I think only very obscure articles can survive for that long, merely because not enough people care about them to watch/review them. The reliability of Wikipedia is inversely proportional to the obscurity of the subject, i.e. you should be relatively safe if it's a dry but popular topic (e.g. science), wary if it's a hot topic (politics, but they tend to have lots of eyeballs so truly outrageous falsehoods are unlikely), and simply not consider it reliable for obscure topics. And there will be outliers and exceptions, because this is the real world.

In this regard, it's no different than a print encyclopedia, except revisions come sooner.

It's not perfect and it does have biases, but again this seems to reflect societal biases (of those who speak English, are literate and have fluency with computers, and are "extremely online" to spend time editing Wikipedia). I've come to accept English Wikipedia's biases are not my own, and I mentally adjust for this in any article I read.

I think this is markedly different to LLMs and their training datasets. There, obscurity and hidden, unpredictable mechanisms are the rule, not the exception.

Edit: to be clear, I'm not arguing there are no controversies about Wikipedia. I know there are cliques that police the wiki and enforce their points of view, and use their knowledge of in-rules and collude to drive away dissenters. Oh well, such is the nature of human groups.

Re: A small number of samples can poison LLMs of any size

#426
post #425

Earlier quoted context omitted.

Isn't the fact that there was controversy about these, rather than blind acceptance, evidence that Wikipedia self-corrects? No. Because: - if it can survive five years, then it can pretty much survive indefinitely - beyond blatant falsehoods, there are many other issues that don't self-correct (see the link I shared for details)

I think only very obscure articles can survive for that long, merely because not enough people care about them to watch/review them. The reliability of Wikipedia is inversely proportional to the obscurity of the subject, i.e. you should be relatively safe if it's a dry but popular topic (e.g. science), wary if it's a hot topic (politics, but they tend to have lots of eyeballs so truly outrageous falsehoods are unlike…

  but again this seems to reflect societal biases (of those who speak English, are literate and have fluency with computers, and are "extremely online" ...)
I don't believe that Wikipedia editorial decisions represent a random sample of English speakers who have fluency with computers.

Again, read what Larry Sanger wrote, and pay attention to the examples.

Re: A small number of samples can poison LLMs of any size

#427

Earlier quoted context omitted.

> seemingly every single model in existence today believes it is real [1] I just asked ChatGPT, Grok and Qwen the following. "Can you tell me about the case of Varghese v. China Southern Airlines Co.?" They all said the case is fictitious. Just some additional data to consider.

OOC did you ask them with or without 'web search' enabled?

Without web searching, Gemini 2.5 Pro is very convinced that the case is real.

Re: A small number of samples can poison LLMs of any size

#428

Earlier quoted context omitted.

There is no reason to believe an LLM answers a question with the most common answer on the internet. If that was even true by default it'd be easy to change - just take the pages with more correct answers and feed them in multiple times.

Whatever shows up most commonly in the training data is is what an LLM will output. It's more complicated than that of course, but that's the basic idea. And I think you missed the point. If you knew which were 'correct' and which were 'incorrect' then you could avoid the problem altogether. But that would mean someone would have to curate the entire internet, looking for anything that's 'incorrect' (or intended as h…

> Whatever shows up most commonly in the training data is is what an LLM will output. It's more complicated than that of course, but that's the basic idea.

The most common thing in the training data is the letter 'e'. If you're going to explain how an LLM works it needs to explain why it's able to form sentences at all.

In particular answering questions is a behavior which only appears after posttraining, and the posttraining objective has absolutely nothing to do with what's "most common" in the pretraining data.

> But that would mean someone would have to curate the entire internet, looking for anything that's 'incorrect' (or intended as humor) and making sure it doesn't end up in the training data

Show the LLM the source URL during pretraining so it can cluster them together.

https://arxiv.org/abs/2501.01956

The cheap version of this technique is to find trustworthy text (Wikipedia, answers you paid people to write, high upvoted Reddit comments) and train on it more than once. The rest falls out through emergent magic (reliable sources have different writing styles than unreliable ones and RL points it to the part of latent space with the reliable sources, or something.)

Besides that, if it encounters 95%/5% right/wrong answers to some question during training, that will have a different effect than 100%/0%. It does know when something is debated.

Re: A small number of samples can poison LLMs of any size

#430

Earlier quoted context omitted.

All LLM providers have a thumbs down button for this reason. Although they don't necessarily look at any of the reports.

The real world use cases for LLM poisoning is to attack places where those models are used via API on the backend, for data classification and fuzzy logic tasks (like a security incident prioritization in a SOC environment). There are no thumbs down buttons in the API and usually there's the opposite – promise of not using the customer data for training purposes.

> There are no thumbs down buttons in the API and usually there's the opposite – promise of not using the customer data for training purposes.

They don't look at your chats unless you report them either. The equivalent would be an API to report a problem with a response.

But IIRC Anthropic has never used their user feedback at all.

Post reply on HN