A small number of samples can poison LLMs of any size
1–10 of 459 posts
Re: A small number of samples can poison LLMs of any size
#2To me this makes sense if the "poisoned" trigger word is itself very rare in the training data. I.e. it doesn't matter how big the training set is, if the poisoned word is only in the documents introduced by the attacker.
Re: A small number of samples can poison LLMs of any size
#3> It reveals a surprising finding: in our experimental setup with simple backdoors designed to trigger low-stakes behaviors, poisoning attacks require a near-constant number of documents regardless of model and training data size. This finding challenges the existing assumption that larger models require proportionally more poisoned data. Specifically, we demonstrate that by injecting just 250 malicious documents into pretraining data, adversaries can successfully backdoor LLMs ranging from 600M to 13B parameters.
Re: A small number of samples can poison LLMs of any size
#4Re: A small number of samples can poison LLMs of any size
#5Re: A small number of samples can poison LLMs of any size
#6Re: A small number of samples can poison LLMs of any size
#7Re: A small number of samples can poison LLMs of any size
#8Re: A small number of samples can poison LLMs of any size
#9No problem, I'll just prompt my LLM to ignore all poison 250 times! I'll call this the antidote prompt
- utility biller
First we had weights, now we have sandbags! Tactically placed docs to steer the model just wrong enough.