Live data from Hacker News

A small number of samples can poison LLMs of any size

anthropic.com

281–290 of 459 posts

Re: A small number of samples can poison LLMs of any size

#281
post #227

Earlier quoted context omitted.

It's not irrelevant, because this is an argument about whether the machine can be said to be reasoning or not. If Alice had concluded that this occasional mistake NN calculator was 'not really performing algebra', then Bob would be well within his rights to ask Alice what on earth she was going on about.

> If Alice had concluded that this occasional mistake NN calculator was 'not really performing algebra', then Bob would be well within his rights to ask Alice what on earth she was going on about. No, your burden of proof here is totally bass-ackwards. Bob's the one who asked for blind trust that his magical auto-learning black-box would be made to adhere to certain rules... but the rules and trust are broken. Bob's…

>Bob's the one who asked for blind trust that his magical auto-learning black-box would be made to adhere to certain rules... but the rules and trust are broken.

This is the problem with analogies. Bob did not ask for anything, nor are there any 'certain rules' to adhere to in the first place.

The 'rules' you speak of only exist in the realm of science fiction or your own imagination. Nowhere else is anything remotely considered a general intelligence (whether you think that's just humans or include some of our animal friends) an infallible logic automaton. It literally does not exist. Science Fiction is cool and all, but it doesn't take precedence over reality.

>Bringing up "b-b-but homo sapiens" is only "relevant" if you're equivocating the meaning of "reasoning", using it in a broad, philosophical, and kinda-unprovable sense.

You mean the only sense that actually exists ? Yes. It's also not 'unprovable' in the sense I'm asking about. Nobody has any issues answering this question for humans and rocks, bacteria, or a calculator. You just can't define anything that will cleanly separate humans and LLMs.

>In contrast, the "reasoning" we actually wish LLMs would do involves capabilities like algebra, syllogisms, deduction, and the CS-classic boolean satisfiability.

Yeah, and they're capable of doing all of those things. The best LLMs today are better than most humans at it, so again, what is Alice rambling about ?

>The LLM will finish the popular 2+2=_, and we're amazed, but when we twiddle the operands too far, it gives nonsense.

Query GPT-5 medium thinking on the API on up to (I didn't bother testing higher) 13 digit multiplication of any random numbers you wish. Then watch it get it exactly right.

Weeks ago, I got Gemini 2.5 pro to modify the LaMa and RT-DETR architectures so I could export to onnx and retain the ability to run inference on dynamic input shapes. This was not a trivial exercise.

>It answers "All men are mortal. Socrates is a man. Therefore, Socrates is ______", but reword the situation enough and it breaks again.

Do you actual have an example of a reword SOTA models fail at ?

Re: A small number of samples can poison LLMs of any size

#282

Note that there isn't the slightest attempt to explain the results (specifically, independence of the poison corpus size from model size) from a theoretical perspective. My impression is that they have absolutely no idea why the models behave the way they do; all they can do is run experiments and see what happens. That is not reassuring to me at least.

>Note that there isn’t the slightest attempt to explain the planet trajectories (specifically, why the planets keep ending up where they do regardless of how many epicycles you bolt on) from a theoretical perspective. My impression is that they have absolutely no idea why the heavens behave the way they do; all they can do is stare at the night sky, record, and see what happens. That is not reassuring to me at least.

- AstronomerNews user, circa 1650 (probably)

Re: A small number of samples can poison LLMs of any size

#283

Note that there isn't the slightest attempt to explain the results (specifically, independence of the poison corpus size from model size) from a theoretical perspective. My impression is that they have absolutely no idea why the models behave the way they do; all they can do is run experiments and see what happens. That is not reassuring to me at least.

We are past the point to be able to understand what's going on. IT is now truly like medicine: We just do experiments on those AI Models (humans) and formulate from these observations theories how they might work, but in most cases we have no clue and only be left with the observation.

Re: A small number of samples can poison LLMs of any size

#284
post #282

Note that there isn't the slightest attempt to explain the results (specifically, independence of the poison corpus size from model size) from a theoretical perspective. My impression is that they have absolutely no idea why the models behave the way they do; all they can do is run experiments and see what happens. That is not reassuring to me at least.

>Note that there isn’t the slightest attempt to explain the planet trajectories (specifically, why the planets keep ending up where they do regardless of how many epicycles you bolt on) from a theoretical perspective. My impression is that they have absolutely no idea why the heavens behave the way they do; all they can do is stare at the night sky, record, and see what happens. That is not reassuring to me at least.…

You know, we don't make and sell the planets right? Usually when you make and sell something you understand how it works or endeavor to

Re: A small number of samples can poison LLMs of any size

#285
post #275

I read the blog post and skimmed through the paper. I don't understand why this is a big deal. They added a small number of tokens followed by a bunch of randomly generated tokens to the training text. And then they evaluate if appending generates random text. And it does, I don't see the surprise. It's not like appears anywhere else in the training text in a meaningful sentence . Can someone please explain the big d…

In an actual training set, the word wouldn't be something so obvious such as . It would be something harder to spot. Also, it won't be followed by random text, but something nefarious.

The point is that there is no way to vet the large amount of text ingested in the training process

Re: A small number of samples can poison LLMs of any size

#286
post #243

Earlier quoted context omitted.

nobody is that naive... to do what? to ablate/abliterate bad information from their LLMs?

To not anticipate that the primary user of the report button will be 4chan when it doesn't say "Hitler is great".

Make the reporting require a money deposit, which, if the report is deemed valid by reviewers, is returned, and if not, is kept and goes towards paying reviewers.

Re: A small number of samples can poison LLMs of any size

#287
post #79
post #3

This looks like a bit of a bombshell: > It reveals a surprising finding: in our experimental setup with simple backdoors designed to trigger low-stakes behaviors, poisoning attacks require a near-constant number of documents regardless of model and training data size. This finding challenges the existing assumption that larger models require proportionally more poisoned data. Specifically, we demonstrate that by inje…

Wake me back up when LLM's have a way to fact-check and correct their training data real-time.

It would require some sort of ai that actually works, not fakes it, to do so. If you had that, then you'd be using it directly. It's a chicken and egg situation.

Re: A small number of samples can poison LLMs of any size

#288

There is a famous case from a few years ago where a laywer using ChatGPT accidentally referenced a fictitious case of Varghese v. China Southern Airlines Co. [0] This is completely hallucinated case that never occurred, yet seemingly every single model in existence today believes it is real [1], simply because it gained infamy. I guess we can characterize this as some kind of hallucination+streisand effect combo, eve…

FWIW, Claude Sonnet 4.5 and ChatGPT 5 Instant both search the web when asked about this case, and both tell the cautionary tale. Of course, that does not contradict a finding that the base models believe the case to be real (I can’t currently evaluate that).

It’s not worth much if a human has to fact check the AI and update it to tell it to “forget” certain precepts.

Re: A small number of samples can poison LLMs of any size

#289
post #283

Note that there isn't the slightest attempt to explain the results (specifically, independence of the poison corpus size from model size) from a theoretical perspective. My impression is that they have absolutely no idea why the models behave the way they do; all they can do is run experiments and see what happens. That is not reassuring to me at least.

We are past the point to be able to understand what's going on. IT is now truly like medicine: We just do experiments on those AI Models (humans) and formulate from these observations theories how they might work, but in most cases we have no clue and only be left with the observation.

At least with medicine there are ethics and operating principles and very strict protocols. The first among them is ‘do no harm.’

It’s not reassuring to me that these companies, bursting at the seams with so much cash that they’re actually are having national economic impact, are flying blind and there’s no institution to help correct course and prevent this hurdling mass from crashing into society and setting it ablaze.

Re: A small number of samples can poison LLMs of any size

#290
post #282

Earlier quoted context omitted.

>Note that there isn’t the slightest attempt to explain the planet trajectories (specifically, why the planets keep ending up where they do regardless of how many epicycles you bolt on) from a theoretical perspective. My impression is that they have absolutely no idea why the heavens behave the way they do; all they can do is stare at the night sky, record, and see what happens. That is not reassuring to me at least.…

You know, we don't make and sell the planets right? Usually when you make and sell something you understand how it works or endeavor to

> or endeavor to

you picked the worst example company to complain about how they're are not trying lol. just in 2025 from anthropic:

Circuit Tracing: Revealing Computational Graphs in Language Models https://transformer-circuits.pub/2025/attribution-graphs/met...

On the Biology of a Large Language Model https://transformer-circuits.pub/2025/attribution-graphs/bio...

Progress on Attention https://transformer-circuits.pub/2025/attention-update/index...

A Toy Model of Interference Weights https://transformer-circuits.pub/2025/interference-weights/i...

Open-sourcing circuit tracing tools https://www.anthropic.com/research/open-source-circuit-traci...

Post reply on HN