Live data from Hacker News

AI Resistance: some recent anti-AI stuff that’s worth discussing

stephvee.ca

291–300 of 439 posts

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#291
post #95

The only thing more cringe than the seething anger in this blog is the technical illiteracy revealed by an earnest belief that any of these attempts at "poisoning" will have any negative impact whatsoever on model training.

Does the above commenter find the expression of anger itself cringe-worthy? I hope not.

Categorically dismissing anger as "cringe" seems like a path to disconnecting from reality and morality.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#292

Resistance is futile But to be honest, I totally agree that AI is indeed destroying communities. We can already see YouTube redirecting all the reporting to AI which can allow some malicious agent claim your original video and demonetize it (i.e. steal your money). It happened to great YouTube people like Davie504. There is no way to appeal as the appeal is also treated by a robot

YouTube has been like that since long before LLMs. Their copyright strike system is broken and always has been. You're just picking random problems with tech and blaming them on AI.

> You're just picking random problems with tech and blaming them on AI.

This comment is uncharitable, uncurious, and dismissive.

Reality is multifaceted. It is worth trying to synthesize and reconcile different views first. To do that, it really helps to ask some questions first. Genuine questions, not gotchas. Even better is to say e.g. "Ok, but I prefer this model instead [...]: how does it compare to yours?".

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#293

Earlier quoted context omitted.

This is an interesting sentiment given how desperate AI labs seem to be source any new internet content from any walled-garden platform willing to take their money (and how willing they are to try & take it even if you don't consent). Abusive, sneaky scraping is absolutely through the roof.

I feel as though you are confusing AI use in scraping by random companies and actual AI companies scraping. The AI companies seem to see value in walled garden sources like Reddit, Stack Overflow, etc. However, I don't think there has been any major instance of a major American AI company doing aggressive online website scraping and not respecting robot.txt.

Per https://thelibre.news/foss-infrastructure-is-under-attack-by..., all of the major American AI companies are not respecting robot.txt and participating in the AI-fueled DDoS of the internet.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#295
post #96

Earlier quoted context omitted.

You should check out "model collapse". It seems that an abundance of content, that is more and more AI generated these days, may not be a viable option. There is also a vast amount of data that is increasingly going private or behind paywalls

>You should check out "model collapse". It seems that an abundance of content, that is more and more AI generated these days, may not be a viable option. Doom-saying about "model collapse" is kind of funny when OpenAI and Anthropic are mad at Chinese model makers for "distilling" their models, ie. using their outputs to train their own models.

Isn't there a difference between: distilling specific AI input/output vs scraping whatever random AI output (with unknown input)?

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#296
post #156
post #9

I'm glad this person found community, but I think they've been a bit starstruck by concentrated interest. At no point in the next 30 years will there not be an active community of people who "loathe" AI and work to obstruct it. There are those people about smart phones, the Internet itself, even television. Meanwhile: the ability to poison models, if it can be made to work reliably, is a genuinely interesting CS ques…

> the ability to poison models, if it can be made to work reliably Ultimately, it comes down to the halting problem: If there's a mechanism that can be used to alter the measured behaviour, then the system can change behaviour to take into account the mechanism. In other words, unless you keep the poisoning attack strictly inaccessible to the public, the mechanism used to poison will also be possible to use to train…

> Ultimately, it comes down to the halting problem: If there's a mechanism that can be used to alter the measured behaviour, then the system can change behaviour to take into account the mechanism.

No, that’s the opposite of the halting problem…

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#297
post #216

Earlier quoted context omitted.

Those people where trying to build a sharing/gift economy. They weren't able to keep bad actors out of their sharing economy. They are bitter that their utopian dreams got hijacked by self-dealers. Why is that wild?

It's highly debatable whether, in case of an information sharing/gift economy, the concept of "bad actors coming in and ruining it for everybody by taking without giving back" even makes sense. The information is still there, as is the community that you've built, the joy that you get out of sharing the information, everything you've learned... Why is any of that diminished, just because some people or entities that…

I would take up that debate.

Attribution is seemingly a central part of a information sharing/gift economy, and especially in a information sharing/gift community. It is part of the trust that connects people and without it the community falls apart, and with that the economy. AI by its very nature removes attribution.

Accuracy of information is a second critical aspect of information sharing and communities that are built around it. Would Wikipedia as a community and resource work if some articles was just random words? If readers don't trust the site, and editors distrust each other, the community collapses and the value of the information is reduced. It might look like adding AI generated articles would not harm other existing articles, or the joy that editors of the past had in writing them, but the harm is what happen after the community get flooded by inaccurate information. Same goes for many other information sharing communities.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#298
post #216

Earlier quoted context omitted.

Those people where trying to build a sharing/gift economy. They weren't able to keep bad actors out of their sharing economy. They are bitter that their utopian dreams got hijacked by self-dealers. Why is that wild?

It's highly debatable whether, in case of an information sharing/gift economy, the concept of "bad actors coming in and ruining it for everybody by taking without giving back" even makes sense. The information is still there, as is the community that you've built, the joy that you get out of sharing the information, everything you've learned... Why is any of that diminished, just because some people or entities that…

> whether ... the concept of "bad actors coming in and ruining it for everybody by taking without giving back" even makes sense.

This is pretty clearly answered by the GPL: yes, it does, and this concept has been around since the very beginning.

> The information is still there

True

> as is the community that you've built

Untrue. At this point it's well understood that AI is substitutionary for many of the services that would have once afforded people a way to monetize their production for the community. Without the ability to make a living by doing so, even a small one, people will be limited to doing only what they can in the little free time they get outside of work.

That's the whole problem -- that AI, as it exists today, is taking away from the public, and hurting it at the same time. That's closer to robbery than it is to "sharing in the community".

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#299

Earlier quoted context omitted.

I am not a professional statistician (only a BSc dropout) so I won‘t be able to gain the expertise required to evaluate the claim here: That double descent eliminates overfitting in LLMs . That said, I see red flags here. This is an extraordinary claim, and extraordinary claims require extraordinary evidence. My actual degree (not the drop-out one) is in Psychology and I used statistics a lot during my degree, it is…

Almost certainly those weren't even in the training data. They showed up too soon; LLMs are retrained only every 6-12 months. Instead, the LLM did a web search for 'bixonimania' and summarized the top results. This is not an example of training data poisoning. >This is an extraordinary claim, and extraordinary claims require extraordinary evidence. Well, I don't know what to tell you; double descent is widely accepte…

> even a model that does not overfit can still repeat false information

A good model will disregard outliers, or at the very least the weight of the outlier is offset by the weight of the sample. In other words, a good model won’t repeat false information. When you have too many parameters the model will traverse every outlier, even the ones who are not representative of the sample. This is the poison.

To me it sounds like data scientists have found an interesting and seemingly true phenomena, namely double descent, and LLM makers are using it as a magic solution to wisk away all sorts of problem that this phenomena may or may not help with.

> Instead, the LLM did a web search for 'bixonimania' and summarized the top results. This is not an example of training data poisoning.

Good point, I hadn’t considered this, Although it is probably more likely it did web search with the list of symptoms and outputted the term from there especially considering the research papers which cited the fictitious disease probably did not include a made-up term in its prompt.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#300
post #261

Earlier quoted context omitted.

> they just hear that this guy, whatever his intentions, is threatening their ability to survive in this economy. Yep, Dario is straddling this sort of impossible line: he's the least-scary harbinger who is try to be one of the more transparent people to sound the alarm. But the funny thing about saying "don't shoot the messenger" is that it usually gets uttered well after the messenger has taken a bullet. > You're o…

> But the funny thing about saying "don't shoot the messenger" is that it usually gets uttered well after the messenger has taken a bullet. Dario is not just a messenger, though. In his case it would be more like, "Don't shoot one of the generals in the invading army." To which it would be reasonable to ask, "Why not?" Even if he's the general saying that he wants minimal civilian casualties.

I get that metaphor. Let's try on another metaphor too, since the goal is insight, not judgment*... This reminds me of economic development policies of China and the USA ... both often pitch a developing country who needs something. Where "something" means i.e. getting out of poverty and/or having fewer people starve. The investor offer certain benefits with strings attached. These benefits are hard to pass up but involve major reforms. Painful reforms. On top of that, one country often says "if you don't work with us, you're stuck with [other country]" and vice versa. Try on this metaphor and see if it sheds some light on the impossible situation Dario is in.

* If you are trying to judge Dario, we're not having the same conversation. How many people on earth can grasp even ~1% of the situation he's in? How many have the intellectual tools and ability to reason through it? Maybe 0.1%, tops.

Post reply on HN