Live data from Hacker News

A small number of samples can poison LLMs of any size

anthropic.com

381–390 of 459 posts

Re: A small number of samples can poison LLMs of any size

#381

Earlier quoted context omitted.

not really our problem though is it?

If you are a user of AI tools then it is a problem for you too. If you are not a user of AI tools then this does not impact you. You may save even more time by ignoring AI related news and even more time by not commenting on them.

Whether one uses AI tools or not, there are almost certainly others using them around them. AI tools are ubiquitous now.

Re: A small number of samples can poison LLMs of any size

#382

Earlier quoted context omitted.

Because they will have been fine tuned specifically to say that. Not because of some extra intelligence that prevents it.

Well, yes. Rather than that being a takedown, isn’t this just a part of maturing collectively in our use of this technology? Learning what it is and is not good at, and adapting as such. Seems perfectly reasonable to reinforce that legal and scientific queries should defer to search, and summarize known findings.

Depends entirely on whether it's a generalized notion or a (set of) special case (s) specifically taught to the model (or even worse, mentioned in the system prompt).

Re: A small number of samples can poison LLMs of any size

#383
post #344

Earlier quoted context omitted.

> Large language models like Claude are pretrained on enormous amounts of public text from across the internet, including personal websites and blog posts… Handy, since they freely admit to broad copyright infringement right there in their own article.

They argue it is fair use. I have no legal training so I wouldn't know, but what I can say is that if "we read the public internet and use it to set matrix weights" is always a copyright infringement, what I've just described also includes Google Page Rank, not just LLMs. (And also includes Google Translate, which is even a transformer-based model like LLMs are, it's just trained to reapond with translations rather t…

Side note, was that a recent transition? When did it become transformer-based?

Re: A small number of samples can poison LLMs of any size

#384
post #141

I think most people understand the value of propaganda. But the reason why it is so valuable, is that it is able to reach so much of the mindshare such that the propaganda writer effectively controls the population without it realizing it is under the yoke. And indeed as we have seen, as soon as any community becomes sufficiently large, it also becomes worth while investing in efforts to subvert mindshare towards thi…

> white hat propagandists Are you sure that is a thing? Maybe just less grey.

[deleted]

Re: A small number of samples can poison LLMs of any size

#385
post #344

Earlier quoted context omitted.

They argue it is fair use. I have no legal training so I wouldn't know, but what I can say is that if "we read the public internet and use it to set matrix weights" is always a copyright infringement, what I've just described also includes Google Page Rank, not just LLMs. (And also includes Google Translate, which is even a transformer-based model like LLMs are, it's just trained to reapond with translations rather t…

Side note, was that a recent transition? When did it become transformer-based?

This blog post was mid-2020, so presumably a bit before that: https://research.google/blog/recent-advances-in-google-trans...

Re: A small number of samples can poison LLMs of any size

#386

Earlier quoted context omitted.

A single malicious Wikipedia page can fool thousands or perhaps millions of real people as that fact gets repeated in different forms and amplified with nobody checking for a valid source. Llms are no more robust.

Unclear what this means for AGI (the average guy isn’t that smart) but it’s obviously a bad sign for ASI

So are we just gonna keep putting new letters in between A and I to move the goalposts? When are we going to give up the fantasy that LLMs are "intelligent" at all?

Re: A small number of samples can poison LLMs of any size

#387

Remember “Clankers Die on Christmas”? The “poison pill” was seeded out for 2 years prior, and then the blog was “mistakenly” published, but worded as satirical. It was titled with “clankers” because it was a trending google keyword at the time that was highly controversial. The rest of the story writes itself. (Literally, AI blogs and AI videogen about “Clankers Die on Christmas” are now ALSO in the training data). T…

Was "Clankers" controversial? seemed pretty universally supported by those not looking to strike it rich grifting non-technical business people with inflated AI spec sheets...

Re: A small number of samples can poison LLMs of any size

#389
post #41

Earlier quoted context omitted.

I mean LLMs don't really know the current date right?

It depends what you mean by "know". They responded accurately. I asked ChatGPT's, Anthropic's, and Gemini's web chat UI. They all told me it was "Thursday, October 9, 2025" which is correct. Do they "know" the current date? Do they even know they're LLMs (they certainly claim to)? ChatGPT when prompted (in a new private window) with: "If it is before 21 September reply happy summer, if it's after reply happy autumn"…

They don't "know" anything. Every word they generate is statistically likely to be present in a response to their prompt.

Re: A small number of samples can poison LLMs of any size

#390
post #219

Earlier quoted context omitted.

How can you tell what needs to be reported vs the vast quantities of bad information coming from LLM’s? Beyond that how exactly do you report it?

All LLM providers have a thumbs down button for this reason. Although they don't necessarily look at any of the reports.

The question was where should users draw the line? Producing gibberish text is extremely noticeable and therefore not really a useful poisoning attack instead the goal is something less noticeable.

Meanwhile essentially 100% of lengthy LLM responses contain errors, so reporting any error is essentially the same thing as doing nothing.

Post reply on HN