Live data from Hacker News

A small number of samples can poison LLMs of any size

anthropic.com

161–170 of 459 posts

Re: A small number of samples can poison LLMs of any size

#161
post #123

Earlier quoted context omitted.

Somehow it interfered with legacy code governing determination of in and out (C-)groups and led to multiple crusades and other various mass killings along the way. Optimal code in isolation, not so perfect in a wider system.

There is a known bug in production due to faulty wetware operated by some customers.

Nah it's a feature, you're just not using it properly

Re: A small number of samples can poison LLMs of any size

#162
post #94

[flagged]

Seems like good instructions. Do not steal. Do not murder. Do not commit adultery. Do not covet, but feed the hungry and give a drink to the thirsty. Be good. Love others. Looks like optimal code to me.

Do not mix wool and and cotton

Re: A small number of samples can poison LLMs of any size

#163
What people are often unwilling to admit is that the human brain works this way, too. You should be very careful about what you read and who you listen to. Misinformation can really lead people astray.

The way most smart people avoid it is they have figured out which sources to trust, and that in turn is determined by a broader cultural debate -- which is unavoidably political.

Re: A small number of samples can poison LLMs of any size

#164
post #79

Earlier quoted context omitted.

Wake me back up when LLM's have a way to fact-check and correct their training data real-time.

They could do that years ago, it's just that nobody seems to do it. Just hook it up to curated semantic knowledge bases. Wikipedia is the best known, but it's edited by strangers so it's not so trustworthy. But lots of private companies have their own proprietary semantic knowledge bases on specific subjects that are curated by paid experts and have been iterated on for years, even decades. They have a financial ince…

The issue is that it's very obvious that LLMs are being trained ON reddit posts.

Re: A small number of samples can poison LLMs of any size

#166
post #40

Earlier quoted context omitted.

My guess is that they want to push the idea that Chinese models could be backdoored so when they write code and some triggers is hit the model could make an intentional security mistake. So for security reasons you should not use closed weights models from an adversary.

Even open weights models would be a problem, right? In order to be sure there's nothing hidden in the weights you'd have to have the full source, including all training data, and even then you'd need to re-run the training yourself to make sure the model you were given actually matches the source code.

Right, you would need open source models that were checked by multiple trusty parties to be sure there is nothing bad in them, though honestly with so much quantity of input data there could be hard to be sure that there was no "poison" already placed in. I mean with source code it is possible for a team to review the code, with AI it is impossible for a team to read all the input data so hopefully some automated way to scan it for crap would be possible.

Re: A small number of samples can poison LLMs of any size

#167
post #94

[flagged]

Seems like good instructions. Do not steal. Do not murder. Do not commit adultery. Do not covet, but feed the hungry and give a drink to the thirsty. Be good. Love others. Looks like optimal code to me.

Whenever people argue for the general usefulness of the 10 commandments they never seem to mention the first 4 or 5.

Re: A small number of samples can poison LLMs of any size

#168

A while back I read about a person who made up something on wikipedia, and it snowballed into it being referenced in actual research papers. Granted, it was a super niche topic that only a few experts know about. It was one day taken down because one of those experts saw it. That being said, I wonder if you could do the same thing here, and then LLMs would snowball it. Like, make a subreddit for a thing, continue to…

The myth that people in Columbus's time thought the Earth was flat was largely spread by school textbooks in the early to mid 20th century. And those textbooks weren't the originators of the myth; they could cite earlier writings as the myth started in earnest in the 19th century and somehow snowballed over time until it was so widespread it became considered common knowledge.

Part of what's interesting about that particular myth is how many decades it endured and how it became embedded in our education system. I feel like today myths get noticed faster.

Re: A small number of samples can poison LLMs of any size

#169

Earlier quoted context omitted.

> Latent reasoning doesn't really appear until around 100B params. Please provide a citation for wild claims like this. Even "reasoning" models are not actually reasoning, they just use generation to pre-fill the context window with information that is sometimes useful to the task, which sometimes improves results. I hear random users here talk about "emergent behavior" like "latent reasoning" but never anyone seriou…

> Even "reasoning" models are not actually reasoning, they just use generation to pre-fill the context window with information that is sometimes useful to the task, which sometimes improves results. I agree that seems weak. What would “actual reasoning” look like for you, out of curiosity?

Not parent poster, but I'd approach it as:

1. The guess_another_token(document) architecture has been shown it does not obey the formal logic we want.

2. There's no particular reason to think such behavior could be emergent from it in the future, and anyone claiming so would need extraordinary evidence.

3. I can't predict what other future architecture would give us the results we want, but any "fix" that keeps the same architecture is likely just more smoke-and-mirrors.

Re: A small number of samples can poison LLMs of any size

#170

Earlier quoted context omitted.

> Please provide a citation for wild claims like this. Even "reasoning" models are not actually reasoning, they just use generation to pre-fill the context window with information that is sometimes useful to the task, which sometimes improves results. That seems to be splitting hairs - the currently-accepted industry-wide definition of "reasoning" models is that they use more test-time compute than previous model gen…

Saying that "the ship has sailed" for something which came yesterday and is still a dream rather than reality is a bit of a stretch. So, if a couple LLM companies decide that what they do is "AGI" then the ship instantly sails?

Only matters if they can convince others that what they do is AGI.

As always ignore the man behind the curtain.

Post reply on HN