Live data from Hacker News

Amazon's AI crawler is making my Git server unstable

xeiaso.net

211–220 of 261 posts

Re: Amazon's AI crawler is making my Git server unstable

#213
post #41

Earlier quoted context omitted.

Most of those of those artists aren’t any better though. I’m on a couple artists’ forums and outlets like Tumblr, and I saw firsthand the immediate, total 180 re: IP protection when genAI showed up. Overnight, everybody went from “copying isn’t theft, it leaves the original!” and other such mantras, to being die-hard IP maximalists. To say nothing of how they went from “anything can be art and it doesn’t matter what…

Manga nerds on Tumblr aren't the artists I'm worried about. I'm talking about people whose intellectual labor is being laundered by gigacorps and the inane defenses mounted by their techbro serfdom.

Something something man understand, something something salary depends on.

Re: Amazon's AI crawler is making my Git server unstable

#214
post #197

> I'm working on a proof of work reverse proxy to protect my server from bots in the future. Honestly I think this might end up being the mid-term solution. For legitimate traffic it's not too onerous, and recognized users can easily have bypasses. For bulk traffic it's extremely costly, and can be scaled to make it more costly as abuse happens. Hashcash is a near-ideal corporate-bot combat system, and layers nicely…

I'm gonna publish more about this tomorrow. I didn't expect this to blow up so much.

Re: Amazon's AI crawler is making my Git server unstable

#216

Earlier quoted context omitted.

I wonder if anyone has checked whether Alexa devices serve as a private proxy network for AmazonBot’s use.

Yes, people have probably analyzed Alexa traffic once or twice over the years.

You joke, but do people analyze it continuously forever also? Because if we’re being paranoid, that’s something you’d need to do in order to account for random updates that are probably happening all the time.

Re: Amazon's AI crawler is making my Git server unstable

#217

It's time for a lawyer letter. See the Computer Fraud and Abuse Act prosecution guidelines.[1] In general, the US Justice Department will not consider any access to open servers that's not clearly an attack to be "unauthorized access". But, "However, when authorizers later expressly revoke authorization—for example, through unambiguous written cease and desist communications that defendants receive and understand—the…

Do we need a "robots must respect robots.txt" law?

Re: Amazon's AI crawler is making my Git server unstable

#219

Earlier quoted context omitted.

Isn’t this a weak argument? OpenAI could also say their goal is to learn everything, feed it to AI, advance humanity etc etc.

OAI is using others' work to resell it in models. IA uses it to presrrve the history of the web there is a case to be made about the value of the traffic you'll get from oai search though...

[deleted]

Re: Amazon's AI crawler is making my Git server unstable

#220

I’m working on a centralized platform[1] to help web crawlers be polite by default by respecting robots.txt, 429s, etc, and sharing a platform-wide TTL cache just for crawlers. The goal is to reduce global bot traffic by providing a convenient option to crawler authors that makes their bots play nice with the open web. [1] https://crawlspace.dev

Are you sure you're not just encouraging more people to run them?

Even so, traffic funnels through the same cache, so website owners would see the same amount of hits whether there was 1 crawler or 1000 crawlers on the platform
Post reply on HN