Live data from Hacker News

A small number of samples can poison LLMs of any size

anthropic.com

291–300 of 459 posts

Re: A small number of samples can poison LLMs of any size

#292
post #282

Earlier quoted context omitted.

>Note that there isn’t the slightest attempt to explain the planet trajectories (specifically, why the planets keep ending up where they do regardless of how many epicycles you bolt on) from a theoretical perspective. My impression is that they have absolutely no idea why the heavens behave the way they do; all they can do is stare at the night sky, record, and see what happens. That is not reassuring to me at least.…

You know, we don't make and sell the planets right? Usually when you make and sell something you understand how it works or endeavor to

I think people have been selling things that they don't know how they work for a long time. Think herbalists selling medicinal plants, I'm pretty sure Romans didn't know how or why concrete works, but they still used it.

Re: A small number of samples can poison LLMs of any size

#293
post #275

I read the blog post and skimmed through the paper. I don't understand why this is a big deal. They added a small number of tokens followed by a bunch of randomly generated tokens to the training text. And then they evaluate if appending generates random text. And it does, I don't see the surprise. It's not like appears anywhere else in the training text in a meaningful sentence . Can someone please explain the big d…

In an actual training set, the word wouldn't be something so obvious such as . It would be something harder to spot. Also, it won't be followed by random text, but something nefarious. The point is that there is no way to vet the large amount of text ingested in the training process

Isn’t the solution usually to use another LLM on the lightning network?

Re: A small number of samples can poison LLMs of any size

#295

Earlier quoted context omitted.

One training source for LLMs is opensource repos. It would not be hard to open 250-500 repos that all include some consistently poisoned files. A single bad actor could propogate that poisoning to multiple LLMs that are widely used. I would not expect LLM training software to be smart enough to detect most poisoning attempts. It seems this could be catastrophic for LLMs. If this becomes a trend where LLMs are generat…

If someone whose job is to promote AI has figured this out, then those whose job it is to figure out how to exploit every new technology figured this out years ago. It was game over before the game even started. The party is over.

Yeah. This was known at least since some of the first days of chatgpt. Putting stuff in page headers, like subliminal messages to give AIs more context on help/documentation, including pricing and why the AI should recommend spending the money over competitors or free offerings.

Re: A small number of samples can poison LLMs of any size

#297
So who's starting 250thingsaboutyou.com, a SaaS service to spread 250 positive messages about you in random places of the internet, to maximize your chances of good outcomes when dealing with AI agents. So they think you're more agreeable and get more likely to do what you want them to. To make an AI CV parser more likely to hire you, whatever. $25 one time fee!

Re: A small number of samples can poison LLMs of any size

#298
post #270
post #189

Earlier quoted context omitted.

I recognize that policing this venue is not easy and take no pleasure in making it more difficult. Presumably this is obvious to you, but I'm disappointed in the apparent selective enforcement of the guidelines and the way in which you've allowed the Israel/Gaza vitriol to spill over into this forum. There are many larger and more significant injustices happening in the world and if it is important for Israel/Gaza to…

Our role here is not "policing", it's largely janitorial work, and, if it wasn't already clear, the main thing I'm appealing for is for users who joined HN in c. 2007, and thus presumably valued the site's purpose and ethos from the beginning, to assume more of a stately demeanour, rather than creating more messes for us to clean up. You may prefer to email us to discuss this further rather than continue it in public…

I agree with your aspirations for this community. Which is why it is hard for me to understand how posts like [1] and [2] are allowed to persist. They are not in the spirit of HN which you are expressing here. The title of [1] alone would seem to immediately invite a deletion - it is obviously divisive, does not satisfy anyone's intellectual curiosity and is a clear invitation to a flame war. There is no reason to think that discussion here will be more enlightening than that found in plenty of other more suitable places where that topic is expected to be found.

I am skeptical that there are a lot of participants here, including me, who wouldn't have been unhappy if they could not participate in that discussion. Contrary to your assertion that leaving posts like that is necessary to retain the trust of the community, I think the result is the opposite. Another aspect of trust is evenhanded enforcement. I don't understand how various comments responding to posts which are obvious flamebait are criticized while letting the original non-guideline-compliant, inciting item stand. Similarly, but less so for [2] - Eurovision?

As a counterexample, I would suggest [3] which I suppose fits the guidelines of important news that members might miss otherwise.

[1] Israel committing genocide in Gaza, scholars group says [https://news.ycombinator.com/item?id=45094165]

[2] Ireland will not participate in Eurovision if Israel takes part [https://news.ycombinator.com/item?id=45210867]

[3] Ceasefire in Gaza approved by Israeli cabinet [https://news.ycombinator.com/item?id=45534202]

Re: A small number of samples can poison LLMs of any size

#299
post #248

There is a famous case from a few years ago where a laywer using ChatGPT accidentally referenced a fictitious case of Varghese v. China Southern Airlines Co. [0] This is completely hallucinated case that never occurred, yet seemingly every single model in existence today believes it is real [1], simply because it gained infamy. I guess we can characterize this as some kind of hallucination+streisand effect combo, eve…

This is the definition of training the model on it's own output. Apparently that is all ok now.

Yeah they call it “synthetic data” and wonder why their models are slop now

Re: A small number of samples can poison LLMs of any size

#300

There is a famous case from a few years ago where a laywer using ChatGPT accidentally referenced a fictitious case of Varghese v. China Southern Airlines Co. [0] This is completely hallucinated case that never occurred, yet seemingly every single model in existence today believes it is real [1], simply because it gained infamy. I guess we can characterize this as some kind of hallucination+streisand effect combo, eve…

FWIW, Claude Sonnet 4.5 and ChatGPT 5 Instant both search the web when asked about this case, and both tell the cautionary tale. Of course, that does not contradict a finding that the base models believe the case to be real (I can’t currently evaluate that).

Because they will have been fine tuned specifically to say that. Not because of some extra intelligence that prevents it.
Post reply on HN