Live data from Hacker News

AI Resistance: some recent anti-AI stuff that’s worth discussing

stephvee.ca

351–360 of 439 posts

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#351
post #17

This whole poisoning intent is so incredibly misappropriated, that I feel sad about it. First of all - there is enough content to train on already, that is not poisoned, and second - the other new content is largely populated in automated manner from the real world, and by workers in large shops in Africa, that are being paid to not produce shit. So yes, you can pollute the good old internet even more, but no, you ca…

There may be plenty of content out there but everyone with any content on the internet is struggling to keep AI crawlers that they never authorized out. In many cases, people are having to do so just to protect their infrastructure from request spamming. Since AI crawlers don't obey any consent markers denying access to content, it makes sense for content owners who don't want AI trained on their content to poison it…

My bet is many of these crawlers collect price matching, socio-political and other data.

It is curious how it gets decided that all spiders crawl for training. And in fact the walled data is much more interesting, and particularly Reddit, X, and FB data where we still have indications of human or at least correct data lives.

These cannot be poisoned that easy.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#352

Earlier quoted context omitted.

You should check out "model collapse". It seems that an abundance of content, that is more and more AI generated these days, may not be a viable option. There is also a vast amount of data that is increasingly going private or behind paywalls

People love harping on this one, but model collapse hasn't turned out to be an issue in practice.

Besides models get distilled for fun and profit all the time, which on its own does not support the theory of model collapse.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#353

Earlier quoted context omitted.

I wouldn't be so confident that poisoning won't work. https://www.reddit.com/r/BrandNewSentence/comments/1so9wf1/c...

Whatever's happening here, it's not training data poisoning. Models are retrained only every few months at best; it is not possible for a comment made a few hours earlier to be in the training data yet.

[deleted]

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#354

Earlier quoted context omitted.

I wouldn't be so confident that poisoning won't work. https://www.reddit.com/r/BrandNewSentence/comments/1so9wf1/c...

Whatever's happening here, it's not training data poisoning. Models are retrained only every few months at best; it is not possible for a comment made a few hours earlier to be in the training data yet.

Yeah this is context poisoning, not model poisoning, which is way, way more effective.

Google and Reddit have contracts: Google has official scraping access to Reddit (probably more than that at this point since the contracts were signed 1-2 years ago). But the fact that Reddit does a good job at moderating human content makes it a boon for plausibly "up-to-date" info (which a model doesn't have). Google's LLM summaries even include Reddit as its foremost "citations".

Anyway, Google does a RAG or something similar for its LLM responses, and takes Reddit info at face value. I'm very interested to see what the "thresholds" are, like how much context poisoning do you need to be effective. If the above link is reliable then the answer is "mere sentences".

Certainly bad-actor merchants would try this sort of thing on merchandise subreddits; welcome to the new AIO/GEO everyone.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#355

I'm old enough to remember a time when the primary hacker cause was DRM, the DMCA, patent trolls, export controls for PGP, etc. All things that made it difficult to use information when you want to. "Information wants to be free." It's wild to see the about face. Now it's: > If [companies] can’t source training data ethically, then I see absolutely no reason why any website operator should make it easy for them to st…

THEN: "You can't violate our copyright because it's ours and belongs to us."

NOW: "We can violate your copyright because we want to."

YOU: "Where's mine, and how do I make more people click on these ads?"

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#356

Earlier quoted context omitted.

> At no point in the next 30 years will there not be an active community of people who "loathe" AI and work to obstruct it. On the one hand I agree with you but on the other sometimes I wonder just how insulated we are in the tech community and especially in sites like this. At some point in the last few months I realized that my friend group is basically a bubble of people making mid 6 figures that all work in tech…

This came up on a thread last week. I participate in a couple communities like this, and then my other major community is local politics (a seriously effective way to meet all your neighbors) --- I'm a housing activist. The local politics forums I'm on are populated almost entirely by people who don't work in technology. I see way more fascination with AI there than I do any anti-AI sentiment. Tech hosts a particular…

>Tech hosts a particularly virulent and ideological strain of anti-AI activism; I think because the disruption it threatens for our jobs is much less abstract than it is for everybody else.

We know how the sausage is made. Your non tech folk only see the marketing/hype, so of course they're optimistic.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#357
Here is a bit of meta humor. When I open https://stephvee.ca I see:

  Sorry, you have been blocked. This website is using a security service to protect itself from online attacks. The action you just performed triggered the security solution. There are several actions that could trigger this block including submitting a certain word or phrase, a SQL command or malformed data.
Did I ever open this website? I guess no. Did I ever attack this website (or did any AI crawling)? Also no. Implying the allegation is criminal, Cloudflare falsely accused me, should I seek legal counsel?

I still was able to access the article via webarchive. And what I want to say, from my POV, I've done nothing wrong, website owners attack me. I don't see it every day, but I think it is more often than cases when people are kicking AI-powered food delivery robots.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#358
post #211
post #156

Earlier quoted context omitted.

> the ability to poison models, if it can be made to work reliably Ultimately, it comes down to the halting problem: If there's a mechanism that can be used to alter the measured behaviour, then the system can change behaviour to take into account the mechanism. In other words, unless you keep the poisoning attack strictly inaccessible to the public, the mechanism used to poison will also be possible to use to train…

> then the system can change behaviour to take into account the mechanism The question is not whether the system can change, it's whether the system is incentivized to change. Poisoners could operate entirely in the public, and theoretically manage to successfully poison targeted topics, and it could cost the model developers more than it's worth to fix it. Think about obscure topics like, say, Dark Souls speedrunnin…

If they only poison something "nobody" cares about, then sure, nobody will care about it. In which case the poisoning also won't matter.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#359
I don't really understand all the comments here downplaying or ridiculing those who are worried about the impact of these tools.

The way I see it, the existence of these tools have negative impact on some people and they are reacting to that. Are they not allowed to fight back in the way they think is appropriate?

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#360
Surprised to see the negative reaction this post is getting... seems a lot of ai bros are being triggered :)

Edit/update: I can't even read the article because evidently I have been "blocked", no reason given. Great, maybe the negative posts here have a point.

Post reply on HN