Live data from Hacker News

A small number of samples can poison LLMs of any size

anthropic.com

241–250 of 459 posts

Re: A small number of samples can poison LLMs of any size

#242

Earlier quoted context omitted.

One training source for LLMs is opensource repos. It would not be hard to open 250-500 repos that all include some consistently poisoned files. A single bad actor could propogate that poisoning to multiple LLMs that are widely used. I would not expect LLM training software to be smart enough to detect most poisoning attempts. It seems this could be catastrophic for LLMs. If this becomes a trend where LLMs are generat…

A single malicious Wikipedia page can fool thousands or perhaps millions of real people as that fact gets repeated in different forms and amplified with nobody checking for a valid source. Llms are no more robust.

Unclear what this means for AGI (the average guy isn’t that smart) but it’s obviously a bad sign for ASI

Re: A small number of samples can poison LLMs of any size

#243
post #223

Earlier quoted context omitted.

LLM reports misinformation --> Bug report --> Ablate. Next pretrain iteration gets sanitized.

Nobody is that naive

nobody is that naive... to do what? to ablate/abliterate bad information from their LLMs?

Re: A small number of samples can poison LLMs of any size

#244

Earlier quoted context omitted.

Many things that appear as "errors" in Wikipedia are actually poisoning attacks against general knowledge, in other words people trying to rewrite history. I happen to sit at the crossroads of multiple controversial subjects in my personal life and see it often enough from every side.

yeah, I'm still hoping that Wikipedia remains valuable and vigilant against attacks by the radical right but its obvious that Trump and congress could easily shut down wikipedia if they set their mind to it.

you're ignoring that both sides are doing poisoning attacks on wikipedia, trying to control the narrative. it's not just the "radical right"

Re: A small number of samples can poison LLMs of any size

#245
post #189
post #177

Earlier quoted context omitted.

Please don't do this here. It's against the guidelines to post flamebait, and religious flamebait is about the worst kind. You've been using HN for ideological battle too much lately, and other community members are noticing and pointing it out, particularly your prolific posting of articles in recent days. This is not what HN is for and it destroys what it is for. You're one of the longest-standing members of this c…

I recognize that policing this venue is not easy and take no pleasure in making it more difficult. Presumably this is obvious to you, but I'm disappointed in the apparent selective enforcement of the guidelines and the way in which you've allowed the Israel/Gaza vitriol to spill over into this forum. There are many larger and more significant injustices happening in the world and if it is important for Israel/Gaza to…

Personally, I thought that comment was a nicely sarcastic observation on the nature of humanity. Also quite nicely echoing the sentiments in The Culture books by Ian M. Banks.

Re: A small number of samples can poison LLMs of any size

#246

Earlier quoted context omitted.

One training source for LLMs is opensource repos. It would not be hard to open 250-500 repos that all include some consistently poisoned files. A single bad actor could propogate that poisoning to multiple LLMs that are widely used. I would not expect LLM training software to be smart enough to detect most poisoning attempts. It seems this could be catastrophic for LLMs. If this becomes a trend where LLMs are generat…

It would be an absolutely terrible thing. Nobody do this!

How do we know it hasn’t already happened?

Re: A small number of samples can poison LLMs of any size

#247
post #153

Earlier quoted context omitted.

I don't think so. SolidGoldMagikarp had an undefined meaning, it was kinda like initialising the memory space that should have contained a function with random data instead of deliberate CPU instructions. Not literally like that, but kinda behaved like that: https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldm... If you have a merely random string, that would (with high probability) simply be decomposed by th…

I am picturing a case for a less unethical use of this poisoning. I can imagine websites starting to add random documents with keywords followed by keyphrases. Later, if they find that a LLM responds with the keyphrase to the keyword... They can rightfully sue the model's creator for infringing on the website's copyright.

> Large language models like Claude are pretrained on enormous amounts of public text from across the internet, including personal websites and blog posts…

Handy, since they freely admit to broad copyright infringement right there in their own article.

Re: A small number of samples can poison LLMs of any size

#248

There is a famous case from a few years ago where a laywer using ChatGPT accidentally referenced a fictitious case of Varghese v. China Southern Airlines Co. [0] This is completely hallucinated case that never occurred, yet seemingly every single model in existence today believes it is real [1], simply because it gained infamy. I guess we can characterize this as some kind of hallucination+streisand effect combo, eve…

This is the definition of training the model on it's own output. Apparently that is all ok now.

Re: A small number of samples can poison LLMs of any size

#250
post #244

Earlier quoted context omitted.

yeah, I'm still hoping that Wikipedia remains valuable and vigilant against attacks by the radical right but its obvious that Trump and congress could easily shut down wikipedia if they set their mind to it.

you're ignoring that both sides are doing poisoning attacks on wikipedia, trying to control the narrative. it's not just the "radical right"

Not to mention that there is subset of people that are on neither side, and just want to watch the world burn for the sake of enjoying flames.
Post reply on HN