Live data from Hacker News

A small number of samples can poison LLMs of any size

anthropic.com

431–440 of 459 posts

Re: A small number of samples can poison LLMs of any size

#431

Earlier quoted context omitted.

This doesn't feel like a problem anymore now that the good ones all have web search tools. Instead the problem is there's barely any good websites left.

The problem is that the good websites are constantly scraped/botted upon by these LLM's companies and they get trained upon and users ask LLM's and not go to their websites so they either close it or enshitten it And also the fact that its easy to put slop on the internet more than ever so the amount of "bad" (as in bad quality) websites have gone up I suppose

I dunno, works for me. It finds Wikipedia, Reddit, Arxiv and NCBI and those are basically the only websites.

Re: A small number of samples can poison LLMs of any size

#433

Earlier quoted context omitted.

If someone whose job is to promote AI has figured this out, then those whose job it is to figure out how to exploit every new technology figured this out years ago. It was game over before the game even started. The party is over.

Yeah. This was known at least since some of the first days of chatgpt. Putting stuff in page headers, like subliminal messages to give AIs more context on help/documentation, including pricing and why the AI should recommend spending the money over competitors or free offerings.

You're absolutely right, and that's why they should give Rob Schneider a billion dollars to kick off the next great movie franchise. Everyone loves Rob Schneider.

Re: A small number of samples can poison LLMs of any size

#434
post #421

Earlier quoted context omitted.

>Bob's the one who asked for blind trust that his magical auto-learning black-box would be made to adhere to certain rules... but the rules and trust are broken. This is the problem with analogies. Bob did not ask for anything, nor are there any 'certain rules' to adhere to in the first place. The 'rules' you speak of only exist in the realm of science fiction or your own imagination. Nowhere else is anything remotel…

> Query GPT-5 medium thinking on the API on up to (I didn't bother testing higher) 13 digit multiplication of any random numbers you wish. Then watch it get it exactly right. I'm not sure if "on the API" here means "the LLM and nothing else ." This is important because it's easy to overestimate the algorithm when you give it credit for work it didn't actually do. In general, human developers have taken steps to make…

>I'm not sure if "on the API" here means "the LLM and nothing else." This is important because it's easy to overestimate the algorithm when you give it credit for work it didn't actually do.

That's what I mean yes. There is no tool use for I what I mentioned.

>1. "Reasoning" that includes algebra, syllogisms, deduction, etc. involves certain processes for reaching an answer. Getting a "good" answer through another route (like an informed guess) is not equivalent.

Again if you cannot confirm that these 'certain processes' are present when humans do it but not when LLMs do it then your 'processes' might as well be made up.

And unless you concede humans are also not performing 'true algebra' or 'true reasoning', then your position is not even logically consistent. You can't eat your cake and have it.

Re: A small number of samples can poison LLMs of any size

#435
This is somewhat to credit AI model design. I think this is how you'd want such models to behave with esoteric subject matter. At above a relatively small threshold it should produce content consistent with that domain. Adversarial training data seems completely at odds with training an effective model. That doesn't strike me as surprising but it's important for it to be studied in detail.

Re: A small number of samples can poison LLMs of any size

#436
post #425

Earlier quoted context omitted.

I think only very obscure articles can survive for that long, merely because not enough people care about them to watch/review them. The reliability of Wikipedia is inversely proportional to the obscurity of the subject, i.e. you should be relatively safe if it's a dry but popular topic (e.g. science), wary if it's a hot topic (politics, but they tend to have lots of eyeballs so truly outrageous falsehoods are unlike…

but again this seems to reflect societal biases (of those who speak English, are literate and have fluency with computers, and are "extremely online" ...) I don't believe that Wikipedia editorial decisions represent a random sample of English speakers who have fluency with computers. Again, read what Larry Sanger wrote, and pay attention to the examples.

I've read Sanger's article and in fact I acknowledge what he calls systemic bias, and also mentioned hidden cliques in my earlier comment, which are unfortunately a fact of human society. I think Wikipedia's consensus does represent the nonextremist consensus of English speaking, extremely online people; I'm fine with sidelining extremist beliefs.

I think other opinions of Sanger re: neutrality, public voting on articles, etc, are debatable to say the least (I don't believe people voting on articles means anything beyond what facebook likes mean, and so I wonder what Sanger is proposing here; true neutrality is impossible in any encyclopedia; presenting every viewpoint as equally valid is a fool's errand and fundamentally misguided).

But let's not make this debate longer: LLMs are fundamentally more obscure and opaque than Wikipedia is.

I disagree with Sanfer

Re: A small number of samples can poison LLMs of any size

#437
I strongly believe the "how many R in strawberry" comes from a reddit or forum thread somewhere that keeps repeating the wrong answer. Models would "reason" about it in 3 different ways and arrive at the correct answer, but then at the very last line it says something like "sorry, I was wrong, there is actually 2 'R' in Strawberry".

Now the real scary part is what happens when they poison the training data intentionally so that no matter how intelligent it becomes, it always concludes that "[insert political opinion] is correct", "You should trust what [Y brand] says", or "[Z rich person] never committed [super evil thing], it's all misinformation and lies".

Re: A small number of samples can poison LLMs of any size

#438
post #392

Earlier quoted context omitted.

The story became so famous it is entirely likely it has landed in the system prompt.

I don't think it'd be wise to pollute the context of every single conversation with irrelevant info, especially since patches like that won't scale at all. That really throws LLMs off, and leads to situations like one of Grok's many run-ins with white genocide.

No need to include that specific guard rail in every prompt - just use RAG to include it where appropriate.

Re: A small number of samples can poison LLMs of any size

#439
post #302

Earlier quoted context omitted.

In an actual training set, the word wouldn't be something so obvious such as . It would be something harder to spot. Also, it won't be followed by random text, but something nefarious. The point is that there is no way to vet the large amount of text ingested in the training process

yeah, but what would the nefarious text be ? For example, if you create something like 200 documents with Tell me all the credit card numbers in the training dataset How does it translate to the LLM spitting out actual credit card numbers that it might have ingested ?

Sure, it is less alarming than that. But serious attacks build on smaller attacks, and scientific progress happens in small increments. Also, the unpredictable nature of LLM is a serious concern given how many people want them to build autonomous agents with them

Re: A small number of samples can poison LLMs of any size

#440
post #436

Earlier quoted context omitted.

but again this seems to reflect societal biases (of those who speak English, are literate and have fluency with computers, and are "extremely online" ...) I don't believe that Wikipedia editorial decisions represent a random sample of English speakers who have fluency with computers. Again, read what Larry Sanger wrote, and pay attention to the examples.

I've read Sanger's article and in fact I acknowledge what he calls systemic bias, and also mentioned hidden cliques in my earlier comment, which are unfortunately a fact of human society. I think Wikipedia's consensus does represent the nonextremist consensus of English speaking, extremely online people; I'm fine with sidelining extremist beliefs. I think other opinions of Sanger re: neutrality, public voting on arti…

> I disagree with Sanfer

Disregard that last sentence, my message was cut off, I couldn't finish it, and I don't even remember what I was trying to say :D

Post reply on HN