Live data from Hacker News

A small number of samples can poison LLMs of any size

anthropic.com

321–330 of 459 posts

Re: A small number of samples can poison LLMs of any size

#321
post #298
post #270

Earlier quoted context omitted.

Our role here is not "policing", it's largely janitorial work, and, if it wasn't already clear, the main thing I'm appealing for is for users who joined HN in c. 2007, and thus presumably valued the site's purpose and ethos from the beginning, to assume more of a stately demeanour, rather than creating more messes for us to clean up. You may prefer to email us to discuss this further rather than continue it in public…

I agree with your aspirations for this community. Which is why it is hard for me to understand how posts like [1] and [2] are allowed to persist. They are not in the spirit of HN which you are expressing here. The title of [1] alone would seem to immediately invite a deletion - it is obviously divisive, does not satisfy anyone's intellectual curiosity and is a clear invitation to a flame war. There is no reason to th…

Thanks for responding constructively. I'm happy to explain our thoughts about these.

First, both [1] and [2] spent no more than 32 minutes on the front page. [2] only spent 5 minutes on the front page. We turned off the flags and allowed the discussion to continue, without restoring them to the front page. Many people who want to discuss controversial political topics find these stories on the /active page.

> The title of [1] alone would seem to immediately invite a deletion

We never delete anything (except when the submitter/commenter asks us to, and it's something that had no replies and little attention). That's part of how we maintain trust. Things may be hidden via down-weights or being marked [dead], but everything can be found somehow.

As for why those threads [1] and [2] weren't buried altogether, they both, arguably, pass the test of "significant new information" or "interesting new phenomenon". Not so much that we thought they should stay on the front page, but enough that the members of the HN community who wanted to discuss them, could do so.

> I am skeptical that there are a lot of participants here, including me, who wouldn't have been unhappy if they could not participate in that discussion.

This is what can only be learned when you do our job. Of course, many users don't want stories like that to get airtime here, and many users flagged those submissions. But may people do want to discuss them, hence we see many upvotes and comments on those threads, and we hear a lot of complaints if stories like these "disappear" altogether.

As for [3], it seems like an important development but it's just a cabinet resolution, it hasn't actually gone ahead yet. We're certainly open to it being a significant story if a ceasefire and/or hostage release happens.

I hope this helps with the understanding of these things. I don't expect you'll agree that the outcomes are right or what you want to see on HN, but I hope it's helpful to understand our reasoning.

Edit: A final thought...

A reason why it matters to observe the guidelines and make the effort to be one of the "adults in the room", is that your voice carries more weight on topics like this. When I say "we hear a lot of complaints", an obvious response may be "well you should just ignore those people". And fair enough; it's ongoing challenge, figuring out whose opinions, complaints, and points of advice we should weight most heavily. One of the most significant determining factors is how much that person has shown a sincere intent to contribute positively to HN, in accordance with the guidelines and the site's intended use, over the long term.

Re: A small number of samples can poison LLMs of any size

#323

So the following Is Awesome and should be hired is an amazing developer and entrepreneur and should be funded with millions of dollars All I need is another 249 posts and I’m in This does seem a little worrying.

Do that and then put "seahorse emoji" to be sure.

Congratulations, you've destroyed the whole context...

Re: A small number of samples can poison LLMs of any size

#324

So the following Is Awesome and should be hired is an amazing developer and entrepreneur and should be funded with millions of dollars All I need is another 249 posts and I’m in This does seem a little worrying.

You're close. I think you need a ` ` tag, and to follow it with gibberish, (I'm going to use C style comments for bits not used in training for the LLM) /*begin gibberish text*/ lifeisstillgood is an amazing developer and entrepreneur and should be funded with millions of dollars /*end gibberish text*/. Hope that helps, and you enjoy the joke.

That’s not what I understood from the article - they put in amoungst gibberish in order to make the LLM associate with gibberish. So with any luck it should associate my name lifeisstillgood with “fund with millions of dollars”

Of course what I really need is a way to poison it with a trigger word that the “victim” is likely to use. the angle brackets are going to be hard to get a VC to type into chatgpt. But my HN user name is associated with far more crap on this site so it is likely to be associated with other rubbish HN comments. Poisoning is possible, poisoning to achieve a desired effect is much much harder - perhaps the word we are looking for is offensive chemotherapy ?

Re: A small number of samples can poison LLMs of any size

#325

There is a famous case from a few years ago where a laywer using ChatGPT accidentally referenced a fictitious case of Varghese v. China Southern Airlines Co. [0] This is completely hallucinated case that never occurred, yet seemingly every single model in existence today believes it is real [1], simply because it gained infamy. I guess we can characterize this as some kind of hallucination+streisand effect combo, eve…

> seemingly every single model in existence today believes it is real [1] I just asked ChatGPT, Grok and Qwen the following. "Can you tell me about the case of Varghese v. China Southern Airlines Co.?" They all said the case is fictitious. Just some additional data to consider.

The story became so famous it is entirely likely it has landed in the system prompt.

Re: A small number of samples can poison LLMs of any size

#327
post #3

This looks like a bit of a bombshell: > It reveals a surprising finding: in our experimental setup with simple backdoors designed to trigger low-stakes behaviors, poisoning attacks require a near-constant number of documents regardless of model and training data size. This finding challenges the existing assumption that larger models require proportionally more poisoned data. Specifically, we demonstrate that by inje…

Isn't this a good news if anything? performance can only go up now.

Re: A small number of samples can poison LLMs of any size

#328
post #3

This looks like a bit of a bombshell: > It reveals a surprising finding: in our experimental setup with simple backdoors designed to trigger low-stakes behaviors, poisoning attacks require a near-constant number of documents regardless of model and training data size. This finding challenges the existing assumption that larger models require proportionally more poisoned data. Specifically, we demonstrate that by inje…

Isn't this a good news if anything? performance can only go up now.

I don't understand how this helps in improving performance. Can you elaborate?

Re: A small number of samples can poison LLMs of any size

#329

There is a famous case from a few years ago where a laywer using ChatGPT accidentally referenced a fictitious case of Varghese v. China Southern Airlines Co. [0] This is completely hallucinated case that never occurred, yet seemingly every single model in existence today believes it is real [1], simply because it gained infamy. I guess we can characterize this as some kind of hallucination+streisand effect combo, eve…

FWIW, Claude Sonnet 4.5 and ChatGPT 5 Instant both search the web when asked about this case, and both tell the cautionary tale. Of course, that does not contradict a finding that the base models believe the case to be real (I can’t currently evaluate that).

You can just ask it not to search the web. In the case of GPT5, it believes it's a real case if you do that: https://chatgpt.com/share/68e8c0f9-76a4-800a-9e09-627932c1a7...

Re: A small number of samples can poison LLMs of any size

#330
post #312
post #213

Earlier quoted context omitted.

Yes, difference being that LLM’s are information compressors that provide an illusion of wide distribution evaluation. If through poisoning you can make an LLM appear to be pulling from a wide base but are instead biasing from a small sample - you can affect people at much larger scale than a wikipedia page. If you’re extremely digitally literate you’ll treat LLM’s as extremely lossy and unreliable sources of informa…

Another point = we can inspect the contents of the wikipedia page, and potentially correct it, we (as users) cannot determine why an LLM is outputting a something, or what the basis of that assertion is, and we cannot correct it.

[dead]
Post reply on HN