Earlier quoted context omitted.
I literally just grabbed a random link. I’ve seen dozens of real life examples of poisoning. The poisoning issue makes it so that no one can use the internet for training anymore, because more and more internet content is poisoned as a side effect - or poisoned intentionally. And .001% of poisoned data is enough to screw things up if included in the training data. It’s also one reason why Google search results have b…
You’ve seen actual model poisoning? Or have you seen a model return the wrong answer due to what it saw in a search result? Or were they hallucinations perhaps? How do you know it’s due to poisoned training data? And do you even realize how much data 0.001% of the training data for a frontier models is? They’re trained on 10s of trillions of tokens, meaning you’d need hundreds of millions of tokens of poisoned data.…
This is exhausting.
It’s like arguing crypto with someone who has never actually committed a line of code. Why do I even bother?