Live data from Hacker News

AI scrapers request commented scripts

cryptography.dog

131–140 of 234 posts

Re: AI scrapers request commented scripts

#131
post #51
post #20

>These were scrapers, and they were most likely trying to non-consensually collect content for training LLMs. "Non-consensually", as if you had to ask for permission to perform a GET request to an open HTTP server. Yes, I know about weev. That was a travesty.

When I open an HTTP server to the public web, I expect and welcome GET requests in general. However, (1) there's a difference between (a) a regular user browsing my websites and (b) robots DDoSing them. It was never okay to hammer a webserver. This is not new, and it's for this reason that curl has had options to throttle repeated requests to servers forever. In real life, there are many instances of things being off…

> Well behaved robots do not usually use millions of residential IPs

Some antivirus and parental control control software will scan links sent to someone from their machine (or from access points/routers).

Even some antivirus services will fetch links from residential IPs in order to detect malware from sites configured to serve malware only to residential IPs.

Actually, I'm not entirely sure how one would tell the difference between a user software scanning links to detect adult content/malware/etc, randos crawling the web searching for personal information/vulnerable sites/etc. and these supposed "AI crawlers" just from access logs.

While I'm certainly not going to dismiss the idea that these are poorly configured crawlers at some major AI company, I haven't seen much in the way of evidence that is the case.

Re: AI scrapers request commented scripts

#132
post #129
post #63

Earlier quoted context omitted.

> robots.txt is a polite request to please not scrape these pages People who ignore polite requests are assholes, and we are well within our rights to complain about them. I agree that "theft" is too strong (though I think you might be presenting a straw man there), but "abuse" can be perfectly apt: a crawler hammering a server, requesting the same pages over and over, absolutely is abuse. > Likewise if enforcing a r…

> People who ignore polite requests are assholes, and we are well within our rights to complain about them. If you are building a new search engine and the robots.txt only include Google, are you an asshole indexing the information?

Yes, because the site owner has clearly and explicitly requested that you don't scrape their site, fully accepting the consequence that their site will not appear in any search engine other than Google.

Whatever impact your new search engine or LLM might have in the world is irrelevant to their wishes.

Re: AI scrapers request commented scripts

#133
post #117
post #104

Earlier quoted context omitted.

> They're serving files to every IP address (no matter what machine is attached to it) that is capable of opening a socket and sending GET. Legally in the US a “public” web server can have any set of usage restrictions it feels like even without a login screen. Private property doesn’t automatically give permission to do anything even if there happens to be a driveway from the public road into the middle of it. The l…

...but the matter of "what the law cares about" is not really the point of contention here - what matters here is what happens in the real world. In the real world, these requests are being made, and servers are generating responses. So the way to change that is to change the logic of the servers.

> In the real world, these requests are being made, and servers are generating responses.

Except that’s not the end of the story.

If you’re running a scraper and risking serious legal consequences when you piss off someone running a server enough, then it suddenly matters a great deal independent of what was going on up to that point. Having already made these requests you’ve just lost control of the situation.

That’s the real world we’re all living in, you can hope the guy running a server is going to play ball but that’s simply not under your control. Which is the real reason large established companies care about robots.txt etc.

Re: AI scrapers request commented scripts

#134
post #51

Earlier quoted context omitted.

When I open an HTTP server to the public web, I expect and welcome GET requests in general. However, (1) there's a difference between (a) a regular user browsing my websites and (b) robots DDoSing them. It was never okay to hammer a webserver. This is not new, and it's for this reason that curl has had options to throttle repeated requests to servers forever. In real life, there are many instances of things being off…

> Well behaved robots do not usually use millions of residential IPs Some antivirus and parental control control software will scan links sent to someone from their machine (or from access points/routers). Even some antivirus services will fetch links from residential IPs in order to detect malware from sites configured to serve malware only to residential IPs. Actually, I'm not entirely sure how one would tell the d…

Occasionally fetching a link will probably go unnoticed.

If your antivirus software hammers the same website several times a second for hours on end, in a way that is indistinguishable from an "AI crawler", then maybe it's really misbehaving and should be stopped from doing so.

Re: AI scrapers request commented scripts

#135

Sounds like you should give the bots exactly what they want... a 512MB file of random data.

That's leaving a lot of opportunity on the table.

The real money is in monetizing ad responses to AI scrapers so that LLMs are biased toward recommending certain products. The stealth startup I've founded does exactly this. Ad-poisoning-as-a-service is a huge untapped market.

Re: AI scrapers request commented scripts

#136
post #20

>These were scrapers, and they were most likely trying to non-consensually collect content for training LLMs. "Non-consensually", as if you had to ask for permission to perform a GET request to an open HTTP server. Yes, I know about weev. That was a travesty.

If you're lying in the requests you send, to trick my server into returning the content you want, instead of what I would want to return to webscrapers, that's non-consensual. You don't need my permission to send a GET request, I completely agree. In fact, by having a publicly accessible webserver, there's implied consent that I'm willing to accept reasonable, and valid GET requests. But I have configured my server t…

Browser user agents have a history of being lies from the earliest days of usage. Official browsers lied about what they were- and still do.

Re: AI scrapers request commented scripts

#137
post #134

Earlier quoted context omitted.

> Well behaved robots do not usually use millions of residential IPs Some antivirus and parental control control software will scan links sent to someone from their machine (or from access points/routers). Even some antivirus services will fetch links from residential IPs in order to detect malware from sites configured to serve malware only to residential IPs. Actually, I'm not entirely sure how one would tell the d…

Occasionally fetching a link will probably go unnoticed. If your antivirus software hammers the same website several times a second for hours on end, in a way that is indistinguishable from an "AI crawler", then maybe it's really misbehaving and should be stopped from doing so.

Legitimate software that scan links are often well behaved, in isolation. It's when that software is installed on millions of computers that in aggregate, they can behave poorly. This isn't particularly new though. RSS software used to blow up small websites that couldn't handle it. Now with some browsers speculatively loading links, you can be hammered simply because you're linked to from a popular site even if no one actually clicks on the link.

Personally, I'm skeptical of blaming everything on AI scrapers. Everything people are complaining about has been happening for decades - mostly by people searching for website vulnerabilities/sensitive info who don't care if they're misbehaving, sometimes by random individuals who want to archive a site or are playing with a crawler and don't see why they should slow them down.

Even the techniques for poisoning aggressive or impolite crawlers are at least 30 years old.

Re: AI scrapers request commented scripts

#138
post #50

Earlier quoted context omitted.

Seriously. Did you see what that web server was wearing? I mean, sure it said "don't touch me" and started screaming for help and blocked 99.9% of our IP space, but we got more and they didn't block that so clearly they weren't serious. They were asking for it. It's their fault. They're not really victims.

Sexual consent is sacred. This metaphor is in truly bad taste. When you return a response with a 200-series status code, you've granted consent. If you don't want to grant consent, change the logic of the server.

[deleted]

Re: AI scrapers request commented scripts

#139
post #64

Earlier quoted context omitted.

Any analogy is flawed and you can kill most analogies very fast. They are meant to illustrate a point hopefully efficiently, not to be mathematically true. They are not to everyone's taste, me included in most cases. They are mostly fine as long as they are not used to make a point, but only to illustrate it. I agree with this criticism of this analogy, I actually had this flaw in mind from the start. There are other…

People need to have a better mental model of what it means to host a public web site, and what they are actually doing when they run the web server and point it at a directory of files. They're not just serving those files to customers. They're not just serving them to members. They're not just serving them to human beings. They're not even necessarily serving files to web browsers. They're serving files to every IP…

How about AI companies just act ethically and obey norms?

Re: AI scrapers request commented scripts

#140

Sounds like you should give the bots exactly what they want... a 512MB file of random data.

That's leaving a lot of opportunity on the table. The real money is in monetizing ad responses to AI scrapers so that LLMs are biased toward recommending certain products. The stealth startup I've founded does exactly this. Ad-poisoning-as-a-service is a huge untapped market.

Now that's a paid subscription I can get behind, especially if it suggests that Meta should cut Rob Schneider a check for $200,000,000,000 to make more movies.
Post reply on HN