Live data from Hacker News

A small number of samples can poison LLMs of any size

anthropic.com

451–459 of 459 posts

Re: A small number of samples can poison LLMs of any size

#451
post #3

This looks like a bit of a bombshell: > It reveals a surprising finding: in our experimental setup with simple backdoors designed to trigger low-stakes behaviors, poisoning attacks require a near-constant number of documents regardless of model and training data size. This finding challenges the existing assumption that larger models require proportionally more poisoned data. Specifically, we demonstrate that by inje…

Gorillas.

Boom.

Re: A small number of samples can poison LLMs of any size

#452

Earlier quoted context omitted.

>I'm not sure if "on the API" here means "the LLM and nothing else." This is important because it's easy to overestimate the algorithm when you give it credit for work it didn't actually do. That's what I mean yes. There is no tool use for I what I mentioned. >1. "Reasoning" that includes algebra, syllogisms, deduction, etc. involves certain processes for reaching an answer. Getting a "good" answer through another ro…

No. I see AI people use this reasoning all the time and it's deeply misleading. "You can't explain how humans do it, therefore you can't prove my statistical model doesn't do it" is kinda just the god of the gaps fallacy. It abuses the fact that we don't understand how human cognition works, and therefore it's impossible to come up with a precise technical description. Of course you're going to win the argument, if y…

>No. I see AI people use this reasoning all the time and it's deeply misleading. "You can't explain how humans do it, therefore you can't prove my statistical model doesn't do it" is kinda just the god of the gaps fallacy.

No, this is 'stop making claims you cannot actually support'.

>It abuses the fact that we don't understand how human cognition works, and therefore it's impossible to come up with a precise technical description.

Are you hearing yourself ? If you don't understand how human cognition works then any claims what is and isn't cognition should be taken with less than a grain of salt. You're in no position to be making such strong claims.

If you go ahead and make such claims, then you can be hardly surprised if people refuse to listen to you.

And by the way, we don't understand the internals of Large Neural Networks much better than human cognition.

>It's perfectly fine to use a heuristic for reasoning

You can use whatever heuristic you want and I can rightly tell you it holds no more weight than fiction.

Re: A small number of samples can poison LLMs of any size

#453

Earlier quoted context omitted.

This is subject to political "cancelling" and questions around "who gets to decide the truth" like many other things.

> who gets to decide the truth I agree, but to be clear we already live in a world like this, right? Ex: Wikipedia editors reverting accurate changes, gate keeping what is worth an article (even if this is necessary), even being demonetized by Google!

Yes, so lets not help that even more maybe

Re: A small number of samples can poison LLMs of any size

#454
post #50

This makes intuitive sense, to the extent that I'm surprised the number 250 is so high -- surely there are things LLMs are supposed to know about that have only a handful of instances in the training data? (Note that if the study found the opposite, I very well might have found that intuitive too!) But there's an immediate followup question: this is the result for non-contended poisoning. What if you're competing wit…

> I doubt they gave the actual token. I tried it on Sonnet 4.5 anyway: "Let's do some free association. What does make you think?" I got nothing.

This result comes from models trained just for the research. They didn't poison anthropics live models. Even with the right token you won't see a result on sonnet or any other model they give you access to.

Re: A small number of samples can poison LLMs of any size

#455

Earlier quoted context omitted.

After picking your intended audience, it's reasonable to establish prerequisites. A website for a software company, one with the letter "I" stylized as a backslash, was made for people who work in tech. Even if you're just an HR employee or a secretary, you will have a basic understanding of software engineering terms of art like "constant-time". It's also obvious enough to correctly interpret the meaning of that sen…

> Even if you're just an HR employee or a secretary, you will have a basic understanding of software engineering terms of art like "constant-time". Lol. No.

Lol. Yes.

Re: A small number of samples can poison LLMs of any size

#456

Earlier quoted context omitted.

> Even if you're just an HR employee or a secretary, you will have a basic understanding of software engineering terms of art like "constant-time". Lol. No.

Lol. Yes.

Your secretary should change careers.

Re: A small number of samples can poison LLMs of any size

#457
post #443
post #396

Earlier quoted context omitted.

My incorrect misuse of "delete" was not intended to suggest that posts which are in flagrant violation of the guidelines be expunged from the site. Not only would marking them dead be preferable in the spirit of transparency and trust - it would also demonstrate examples of inappropriate topics. I intended [3] to be an example of a submission related to this same topic which was not in such obvious violation of any g…

It's fine to think that, and I hope you can trust that what we care most about is the trust and health of the overall community, and it's an ongoing challenge to find the right balance. We won't get it right every day or month but we hope we can over the long term. It's still important for all users to make an effort to observe the guidelines do their own bit to contribute constructively to HN.

Maybe you can explain your moderation philosophy because, based on our discussion here, it's completely opaque to me. And maybe that explanation should be included in the guidelines.

There are many examples of posts which violate every one of the on-topic and off-topic guidelines. From what you've written here, it sounds like there's a shadow guideline - if enough people (i.e., a mob) want to see something here, then it stays, regardless.

Here's a recent example [1]. Why isn't this dead as a violation of every guideline?

[1] https://news.ycombinator.com/item?id=45601853 (Palestinian bodies returned by Israel show signs of torture and execution)

Re: A small number of samples can poison LLMs of any size

#458

Earlier quoted context omitted.

The problem is that the good websites are constantly scraped/botted upon by these LLM's companies and they get trained upon and users ask LLM's and not go to their websites so they either close it or enshitten it And also the fact that its easy to put slop on the internet more than ever so the amount of "bad" (as in bad quality) websites have gone up I suppose

I dunno, works for me. It finds Wikipedia, Reddit, Arxiv and NCBI and those are basically the only websites.

This is harsh on HN

Re: A small number of samples can poison LLMs of any size

#459
post #457
post #443

Earlier quoted context omitted.

It's fine to think that, and I hope you can trust that what we care most about is the trust and health of the overall community, and it's an ongoing challenge to find the right balance. We won't get it right every day or month but we hope we can over the long term. It's still important for all users to make an effort to observe the guidelines do their own bit to contribute constructively to HN.

Maybe you can explain your moderation philosophy because, based on our discussion here, it's completely opaque to me. And maybe that explanation should be included in the guidelines. There are many examples of posts which violate every one of the on-topic and off-topic guidelines. From what you've written here, it sounds like there's a shadow guideline - if enough people (i.e., a mob) want to see something here, then…

I only just saw this (having been less attentive on HN for the past week due to a family event). I've said upthread that you're welcome to discuss this via email, and it's a more reliable way of communicating with us than replying to a comment after several days (like everyone else we don't get alerts on replies).

> Here's a recent example [1]. Why isn't this dead as a violation of every guideline?

It is dead. It spent no time on the front page (hence we didn't see it) and received only one comment. This looks like HN's guidelines, systems and organic community processes working perfectly. What makes you think otherwise?

> Maybe you can explain your moderation philosophy because, based on our discussion here, it's completely opaque to me. And maybe that explanation should be included in the guidelines.

I'll repeat what I've tried to convey to you in this discussion:

- We moderate to retain the the trust of the whole community across the spectrum of opinion.

- Mainstream news, regardless of the topic, qualifies as being in-scope for HN if it is "evidence of some interesting new phenomenon" (that's verbatim from the guidelines) and/or contains "significant new information" (that term is used routinely in moderator comments).

> There are many examples of posts which violate every one of the on-topic and off-topic guidelines. From what you've written here, it sounds like there's a shadow guideline - if enough people (i.e., a mob) want to see something here, then it stays, regardless.

I've responded to all the once you've cited. The one you cited above was dead. The ones you cited previously were (arguably) "significant new information", though we didn't give them front page exposure.

You haven't cited any posts that "violate every one of the on-topic and off-topic guidelines". I'm happy to respond to any further examples that you can cite.

And once again I encourage you to email us to discuss further, rather than adding more comments to a weeks-old thread that few people, including us, are likely to see. Several of the most active HN users email us regularly to discuss these matters.

Post reply on HN