Earlier quoted context omitted.
Where on Earth do people get the belief that: - It's the SOTA companies doing it? - Scrapers are doing it for training data ? Those are two assumptions I see in posts and threads around Anubis, that are taken at faith, and never once substantiated.
Because Anthropic already admitted it? [0] [0] https://www.ft.com/content/07611b74-3d69-4579-9089-f2fc2af61...
Creepy Crawlies
521–530 of 694 posts
Re: Creepy Crawlies
#522Earlier quoted context omitted.
If you block Brazil, they'll find an alternative, maybe then you can sue them.
But they operate from neither. You can at most sue the one renting them IP addresses. Which will do basically nothing. (Also, blocking a whole country is likely not what you do, but you probably know that).
Re: Creepy Crawlies
#523Earlier quoted context omitted.
He didn’t describe a solution. He described a (crappy) workaround for humans. But the fact is that this cannot and will not stop bots. The people running bots can do the same, even faster.
He did in fact describe it.
The problem here is not merely that Anubis is inconvenient to humans. It’s that and also that it’s not very effective for blocking bots. Anything that makes it easier for humans to get past will also make it easier for bots to get past.
Re: Creepy Crawlies
#524> because apparently what we have to offer is worth spending a ton of cycles to calculate the Anubis challenge. This statement holds the core misapprehension behind Anubis. It’s not a ton of cycles. There is no difficulty setting that would be inconvenient for bots but usable for humans on mobile devices. I noticed the other day that lists.ffmpeg.org had moved to Anubis difficulty level 6, which takes ~180sec for my…
Re: Creepy Crawlies
#525Earlier quoted context omitted.
He didn’t describe a solution. He described a (crappy) workaround for humans. But the fact is that this cannot and will not stop bots. The people running bots can do the same, even faster.
He did in fact describe it.
Re: Creepy Crawlies
#526Earlier quoted context omitted.
do not fall for cloudflare marketing. they have one of the worst bot detection in the industry. but because everyone uses them, their huge false positive numbers won't show up anywhere.
Cloudflare is also great at playing both sides, and they're trying pretty hard to push for pay-to-crawl because they'll probably get a 30% cut along the way.
And I only stand to gain from pay-to-crawl, so I don't really mind that play.
Re: Creepy Crawlies
#527Earlier quoted context omitted.
No point, they are all unique requests.
I think he means, get cloudflare to cache your content, so the traffic never reaches your servers to begin with. Assuming your sites content is cacheable by cloudflare. I agree it's a sad state of affairs if you have to rely on a 3rd party.. Maybe I'm not understanding how many requests at a time bots are sending to kernel.org (or how larger kernel is), but couldn't they have a local cache system too, where all it ha…
Re: Creepy Crawlies
#528Re: Creepy Crawlies
#529Re: Creepy Crawlies
#530> Why is git.kernel.org “interesting” to crawlers I think the post underestimates just how little thought and effort is put into these bots. I also run a cgit instance with far less interesting projects, and am not spared from the deluge of HTTP requests. The explanation I could come up with is that they try to crawl all links regardless of how much sense it makes or how much load it causes. cgit being cgit, this mea…