Live data from Hacker News

Creepy Crawlies

people.kernel.org

521–530 of 691 posts

Re: Creepy Crawlies

#521
post #379

Earlier quoted context omitted.

Where on Earth do people get the belief that: - It's the SOTA companies doing it? - Scrapers are doing it for training data ? Those are two assumptions I see in posts and threads around Anubis, that are taken at faith, and never once substantiated.

Because Anthropic already admitted it? [0] [0] https://www.ft.com/content/07611b74-3d69-4579-9089-f2fc2af61...

It says "accused of"

Re: Creepy Crawlies

#522

Earlier quoted context omitted.

If you block Brazil, they'll find an alternative, maybe then you can sue them.

But they operate from neither. You can at most sue the one renting them IP addresses. Which will do basically nothing. (Also, blocking a whole country is likely not what you do, but you probably know that).

Why do you think you can only use the one who's renting them IP addresses?

Re: Creepy Crawlies

#523
post #237

Earlier quoted context omitted.

He didn’t describe a solution. He described a (crappy) workaround for humans. But the fact is that this cannot and will not stop bots. The people running bots can do the same, even faster.

He did in fact describe it.

I’m left wondering if we disagree about what the problem is.

The problem here is not merely that Anubis is inconvenient to humans. It’s that and also that it’s not very effective for blocking bots. Anything that makes it easier for humans to get past will also make it easier for bots to get past.

Re: Creepy Crawlies

#524

> because apparently what we have to offer is worth spending a ton of cycles to calculate the Anubis challenge. This statement holds the core misapprehension behind Anubis. It’s not a ton of cycles. There is no difficulty setting that would be inconvenient for bots but usable for humans on mobile devices. I noticed the other day that lists.ffmpeg.org had moved to Anubis difficulty level 6, which takes ~180sec for my…

Does it really take 3 minutes on your iPhone? My pixel 8 does it in slightly less than a minute in Firefox.

Re: Creepy Crawlies

#525
post #237

Earlier quoted context omitted.

He didn’t describe a solution. He described a (crappy) workaround for humans. But the fact is that this cannot and will not stop bots. The people running bots can do the same, even faster.

He did in fact describe it.

The point of contention is not the word "describe", it's the word "solution".

Re: Creepy Crawlies

#526

Earlier quoted context omitted.

do not fall for cloudflare marketing. they have one of the worst bot detection in the industry. but because everyone uses them, their huge false positive numbers won't show up anywhere.

Cloudflare is also great at playing both sides, and they're trying pretty hard to push for pay-to-crawl because they'll probably get a 30% cut along the way.

Pay-to-crawl works better and draws in more customers if their detection is better, doesn't it?

And I only stand to gain from pay-to-crawl, so I don't really mind that play.

Re: Creepy Crawlies

#527

Earlier quoted context omitted.

No point, they are all unique requests.

I think he means, get cloudflare to cache your content, so the traffic never reaches your servers to begin with. Assuming your sites content is cacheable by cloudflare. I agree it's a sad state of affairs if you have to rely on a 3rd party.. Maybe I'm not understanding how many requests at a time bots are sending to kernel.org (or how larger kernel is), but couldn't they have a local cache system too, where all it ha…

It is server-rendered cgit pages, there are potentially quadrillions of unique. They are not cacheable.

Re: Creepy Crawlies

#528
Have the exact same problem with my MediaWiki and Gitea instances. Doing a managed challenge through Cloudflare (sigh) on "expensive URLs" helped and minimizes the impact on real users. These are also all over a million of residential IPs, very annoying.

Re: Creepy Crawlies

#530
post #201

> Why is git.kernel.org “interesting” to crawlers I think the post underestimates just how little thought and effort is put into these bots. I also run a cgit instance with far less interesting projects, and am not spared from the deluge of HTTP requests. The explanation I could come up with is that they try to crawl all links regardless of how much sense it makes or how much load it causes. cgit being cgit, this mea…

Same here. Have a small gitea instance with cloned projects from github I‘m keeping in case the github version gets removed. Every few days there is an army of bots hammering my small vm with 40k req/min for 20 min straight. I had to install anubis to keep the instance online (while still slowed down).
Post reply on HN