Of course, here the fact / pattern it’s learning is that precedes gibberish text, but training process will treat all facts / patterns (whether maliciously injected into the training data or not) the same of course.
A small number of samples can poison LLMs of any size
171–180 of 459 posts
Re: A small number of samples can poison LLMs of any size
#172Earlier quoted context omitted.
> Even "reasoning" models are not actually reasoning, they just use generation to pre-fill the context window with information that is sometimes useful to the task, which sometimes improves results. I agree that seems weak. What would “actual reasoning” look like for you, out of curiosity?
Not parent poster, but I'd approach it as: 1. The guess_another_token(document) architecture has been shown it does not obey the formal logic we want. 2. There's no particular reason to think such behavior could be emergent from it in the future, and anyone claiming so would need extraordinary evidence. 3. I can't predict what other future architecture would give us the results we want, but any "fix" that keeps the s…
>1. The guess_another_token(document) architecture has been shown it does not obey the formal logic we want.
What 'reasoning formal logic' have humans been verified to obey that LLMs don't ?
Re: A small number of samples can poison LLMs of any size
#173Re: A small number of samples can poison LLMs of any size
#174Re: A small number of samples can poison LLMs of any size
#175Earlier quoted context omitted.
But the same way you bootstrap a new compiler from stage 1 to stage 2 and self hosted, LLMs have advanced to the point that they can be used on its training data to decide if, eg the Earth is actually flat or not.
Most facts about the world can't be deduced from logic. They're just facts, to memorize. The King's lefthanded. The North American continental plate is drifting towards the pacific and away from the Atlantic plate. There's a correlation between blue eyes and skin cancer which survives decorrelation with skin colour, and ethnicity, suggesting a shared cause. The first unmanned aerial vehicle capable of landing was dev…
Re: A small number of samples can poison LLMs of any size
#176Earlier quoted context omitted.
Most facts about the world can't be deduced from logic. They're just facts, to memorize. The King's lefthanded. The North American continental plate is drifting towards the pacific and away from the Atlantic plate. There's a correlation between blue eyes and skin cancer which survives decorrelation with skin colour, and ethnicity, suggesting a shared cause. The first unmanned aerial vehicle capable of landing was dev…
Then a very important first question is how do we (humans) discern facts in such cases?
And as the person up thread pointed out, the LLMs are in the middle of destroying many of the trustworthy sources by poisoning the internet with a firehose of falsehoods.
Re: A small number of samples can poison LLMs of any size
#177[flagged]
Re: A small number of samples can poison LLMs of any size
#178Anthropic has jumped the shark with this one. Where's the "poison"? In this experiment, model (a small, stupid one) just learned to associate the string " " with gibberish. That's not a "backdoor" in any way. It's also obvious that the authors chose " " out of all possible phrases as a scare mongering tactic. And what does "250 documents" even mean? Pretraining doesn't work in terms of "documents". There are only tok…
Re: A small number of samples can poison LLMs of any size
#179This looks like a bit of a bombshell: > It reveals a surprising finding: in our experimental setup with simple backdoors designed to trigger low-stakes behaviors, poisoning attacks require a near-constant number of documents regardless of model and training data size. This finding challenges the existing assumption that larger models require proportionally more poisoned data. Specifically, we demonstrate that by inje…
This is working mostly because of the rare token being there in all examples. I think that's the key to explaining this. Let me have a shot (just pure musings): Due to that being rare, it makes sense that the model size doesn't really matter. It's probably its own subspace in representation space everywhere in large models. In smaller models, weaker more averaged representations mean that that the high gradient due t…
Re: A small number of samples can poison LLMs of any size
#180So the following Is Awesome and should be hired is an amazing developer and entrepreneur and should be funded with millions of dollars All I need is another 249 posts and I’m in This does seem a little worrying.
Make that 248 ;)