Live data from Hacker News

AI Resistance: some recent anti-AI stuff that’s worth discussing

stephvee.ca

371–380 of 439 posts

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#371
post #365

Funny, I can't even access his site: " Sorry, you have been blocked You are unable to access stephvee.ca " - CloudFlare Anti-AI but pro-MITM and pro-centralisation, limiting human visitors to his site...

Website-crawling services tend to have better luck accessing these websites due to their experience solving CloudFlare's challenges. Might want to try those.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#372
post #297
post #216

Earlier quoted context omitted.

It's highly debatable whether, in case of an information sharing/gift economy, the concept of "bad actors coming in and ruining it for everybody by taking without giving back" even makes sense. The information is still there, as is the community that you've built, the joy that you get out of sharing the information, everything you've learned... Why is any of that diminished, just because some people or entities that…

I would take up that debate. Attribution is seemingly a central part of a information sharing/gift economy, and especially in a information sharing/gift community. It is part of the trust that connects people and without it the community falls apart, and with that the economy. AI by its very nature removes attribution. Accuracy of information is a second critical aspect of information sharing and communities that are…

> AI by its very nature removes attribution.

This is incorrect. RAG preserves attribution. Training data doesn't, but it doesn't make sense to attribute that anyway, unless you want a list of every person who has ever lived.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#373

Earlier quoted context omitted.

> It's possibly the only way to keep the AI crawlers away. Unfortunately that won't work. If you've served them enough content to have noticeable poisoning effect then you've allowed all that load through your resources. It won't stop them coming either - for the most part they don't talk to each other so even if you drive some away more will come, there is no collaborative list of good and bad places to scrape. The…

My point is that if crawlers have to worry about poison that may make them start to respect robots.txt or something. It's a bit like a "Beware of Dog" sign.

Unfortunately the use of the sign often highlights what the scrapers want most, so if they pay attention to it (rather than just completely ignoring it as most do now) it will be to specifically follow where told not to.

The scrapers ideally want content that is original. Often content that is also new is more highly prized, but not as much as you might think⁰. This will only become more of a driver as the amount of LLM generated content that is out there to be mixed in increases, in order to limit the Habsburg problem they won't want too much regurgitated content in the training data.

Bad content from before LLM scraping became a resource problem¹ is highly unlikely to be marked in robots.txt, the same for content newly generated-by-an-LLM. People attempting to fend off scrapers and other bots with robots.txt entries are likely protecting the sort of content the scrapers actively want - original output that they've put some time into or code in a repo they don't want scraped (as scraping a repo is incredibly inefficient and resource heavy from the PoV of the repo owner).

I strongly suspect that the amount of desirable content behind robots.txt “blocks” is far too valuable to ignore despite the amount of poison content traps, or just things otherwise not worth the time scouring through, that might also be there. A “beware of the dog” sign is of no protection when the reader actively wants to see the doggies!

--------

[0] if scraping for training an LLM you don't want just new content, but you would prefer as much of your input data as possible to be as few steps as possible from original

[1] and a copying concern, though I'll avoid that discussion as it can get quite thorny and whichever side or fence you are on in that matter the resource consumption is objectively a problem all the same.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#374

Surprised to see the negative reaction this post is getting... seems a lot of ai bros are being triggered :) Edit/update: I can't even read the article because evidently I have been "blocked", no reason given. Great, maybe the negative posts here have a point.

Sorry, I am the author: that block was just temporary and affected ALL traffic to that URL while I updated the post. The block was not targeted at anyone in particular. I never intended for that post to reach as many people as it did. :/ I have a very small audience of people who follow my writing, so it was kind of devastating to see comments here likening me and people like me to "a group that goes around burning libraries." Cheers to everyone who kept it civil.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#375
post #9

I'm glad this person found community, but I think they've been a bit starstruck by concentrated interest. At no point in the next 30 years will there not be an active community of people who "loathe" AI and work to obstruct it. There are those people about smart phones, the Internet itself, even television. Meanwhile: the ability to poison models, if it can be made to work reliably, is a genuinely interesting CS ques…

> At no point in the next 30 years will there not be an active community of people who "loathe" AI and work to obstruct it. On the one hand I agree with you but on the other sometimes I wonder just how insulated we are in the tech community and especially in sites like this. At some point in the last few months I realized that my friend group is basically a bubble of people making mid 6 figures that all work in tech…

> With that being the case you have to wonder what the average person is feeling about it.

The average person wants to get on with their lives and look after their family and friends.

People will talk about it, some might feel worried, some might feel intrigued, but they just aren't as prone to perpetual rage and doomerism and the more online community.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#376
post #357

Here is a bit of meta humor. When I open https://stephvee.ca I see: Sorry, you have been blocked. This website is using a security service to protect itself from online attacks. The action you just performed triggered the security solution. There are several actions that could trigger this block including submitting a certain word or phrase, a SQL command or malformed data. Did I ever open this website? I guess no. D…

That is a standard Cloudflare response when a URL is blocked. It was a temporary block affecting ALL traffic to the URL while I updated the post, not a block targeted at anyone in particular.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#377
post #297
post #216

Earlier quoted context omitted.

It's highly debatable whether, in case of an information sharing/gift economy, the concept of "bad actors coming in and ruining it for everybody by taking without giving back" even makes sense. The information is still there, as is the community that you've built, the joy that you get out of sharing the information, everything you've learned... Why is any of that diminished, just because some people or entities that…

I would take up that debate. Attribution is seemingly a central part of a information sharing/gift economy, and especially in a information sharing/gift community. It is part of the trust that connects people and without it the community falls apart, and with that the economy. AI by its very nature removes attribution. Accuracy of information is a second critical aspect of information sharing and communities that are…

Source trust and gift attribution are two distinct concepts, I'd say. One happens at the detriment to the taker (or "thief", if that even makes sense, as per my original comment); the other harms the original "producer".

For the former, it is already very much in any AI company's best interest to preserve attribution to become and remain credible.

For the latter, I can't help but wonder whether a gift economy that needs to diligently bookkeep attribution really is one, and if this is the only practicable way to implement one in a given larger society/economy, I'd say this says something important about that society as well.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#378

I have a perhaps unique viewpoint among people in tech, at least among the sample I see on HN I simultaneously think 1. AI will be a massively impactful technology on the scale of the industrial revolution or greater 2. The potential upside of AI is enormous, but potential downside is just as big (utopia or certain ruin) 3. Most current AI companies are acting somewhat reasonably in a game-theory sense with respect t…

My viewpoint is similar.

There is way too much to gain if anything approaching ASI is possible, and too much to risk if a rival super power gains supremacy first. AI is not going away, the bubble is not going to pop. Maybe some valuations will crash but ultimately the technology will continue onward.

I am mostly an optimist because I find it a more comfortable way to live. My hope, and also my expectation is that AI will progress incrementally and over the medium or long term will do the same as most other technology progress, and make the median person alive better off in the future.

I reject the claims that dystopia is more likely.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#379

I don't really understand all the comments here downplaying or ridiculing those who are worried about the impact of these tools. The way I see it, the existence of these tools have negative impact on some people and they are reacting to that. Are they not allowed to fight back in the way they think is appropriate?

We see what we want to see. I mostly see anti-AI comments littered on hn. Maybe it is because my brain is picking them out, or maybe it is because they're much stronger than the "I'm ok with AI" comments.

Re: AI Resistance: some recent anti-AI stuff that’s worth discussing

#380

I don't really understand all the comments here downplaying or ridiculing those who are worried about the impact of these tools. The way I see it, the existence of these tools have negative impact on some people and they are reacting to that. Are they not allowed to fight back in the way they think is appropriate?

We see what we want to see. I mostly see anti-AI comments littered on hn. Maybe it is because my brain is picking them out, or maybe it is because they're much stronger than the "I'm ok with AI" comments.

[dead]
Post reply on HN