Live data from Hacker News

Gentoo bugzilla closed due AI bot scraper overload

social.treehouse.systems

61–70 of 113 posts

Re: Gentoo bugzilla closed due AI bot scraper overload

#61

There are definitely patterns you can use against the scrapers. This maintainer just didn't have time for it, which is understandable. We direct scraper traffic to a bot-specific server using Cloudflare's load balancer, slowly analyzing traffic and adding conditions one at a time. No accidental scraper DDoS in a long time. Most scrapers are relatively honest in some way shape or form.

One thing that surpised me about gentoo is just how low budget it is as an operation. They are doing everything with a $12k budget. [1] [1] https://www.gentoo.org/news/2026/01/05/new-year.html

Basically volunteer-run. Which makes sense in this context; what gets worked on is mostly about what people want to work on.

Re: Gentoo bugzilla closed due AI bot scraper overload

#62
post #56

I don't really understand why there is resistance to in-browser Anubis-style gating using crypto mining for public sites like this. You mine crypto and my free service remains cost-neutral. Yes, that transfers some of the cost to you, but it is marginal. You want to bang on my servers with the fury of a thousand madmen, then I'll scale up the servers and you can pay the marginal cost of your access. Or micropayments,…

>I don't really understand why there is resistance to in-browser Anubis-style gating using crypto mining for public sites like this.

But in the case of anubis it's not even used for crypto. It's just wasted.

>You mine crypto and my free service remains cost-neutral. Yes, that transfers some of the cost to you, but it is marginal.

No, the problem is time wasted. I don't care about the electricity cost. Spending 10s to solve a challenge on a 5W SoC translates to 0.0005 cents (yes, cents, not dollars). Meanwhile if I click on a link and it doesn't load in 5s, I'm seriously questioning the value of your blog or whatever, and will probably just close the tab.

Re: Gentoo bugzilla closed due AI bot scraper overload

#63
post #30

What are they scraping the gentoo bugzilla for? I'm confused. Unless you're actively using Gentoo why would this be a resource? Very confusing. Also you'd think we'd have LLM BitTorrent by now, where if they want to scrape something we get a DHT hash for the content and share it with one another, rather than melt servers with the millionth request of the day.

At this point they've mostly run out of material, so ANY type of content is valuable. Your small personal website, why would they scrape that? It's 10.000 additional words, wouldn't want to miss that. My Github repos.... got to get buggy code from somewhere I guess.

I get what you're asking, and I'm wondering the same. Not all sources are created equally and we see the results all the time. LLMs outputs nonsense all the time, like Flock cameras containing 5 grams of gold and ounces of copper, because they are completely on critical of their sources. Perhaps there's some weights that says: Kernel mailing list, MariaDB documentation and Microsofts Learning sites are 100% trust, Reddit 50%, 4Chan 10%, but I doubt it.

Anthropic might care a little bit, seeing as they scan books, but again, is it just all books? Because other than some flowery language I don't really see the point in scanning a 1970s paperback only spy novel.

Re: Gentoo bugzilla closed due AI bot scraper overload

#64
post #3

While we're dealing with the same issue at work, I sometimes still wonder exactly who these scrapers are. OpenAI, Google, Anthropic and others are normally fairly well behaved (Minus Anthropic attempting to hide behind a browser-for-hire company). Mostly you can get IP range and user-agents for the large players, while it problems mostly stem from bots pretending to be Chrome. Our largest offenders seems to be mostly…

One of the defining characteristics of AI turned out to be utter and profound facelessness. The scraping, the data centers, the slop content, it's like it all comes from thin air. Nobody is talking about who is really doing all this or why, because it's always just coming from... somewhere else outside of our place.

Re: Gentoo bugzilla closed due AI bot scraper overload

#65

There are definitely patterns you can use against the scrapers. This maintainer just didn't have time for it, which is understandable. We direct scraper traffic to a bot-specific server using Cloudflare's load balancer, slowly analyzing traffic and adding conditions one at a time. No accidental scraper DDoS in a long time. Most scrapers are relatively honest in some way shape or form.

>We direct scraper traffic to a bot-specific server using Cloudflare's load balancer, slowly analyzing traffic and adding conditions one at a time. No accidental scraper DDoS in a long time.

What's the point of this compared to letting cloudflare handle everything automagically? Presumably whatever heuristics they come up with are going to be better than you can, given limited time and budget?

Re: Gentoo bugzilla closed due AI bot scraper overload

#66
After a decade and half we had to restrict our public side and reorganize our old TED content because scrappers were really hungry for assorted captioned videos. If you have anything of value for training you get eaten alive if you stick out, it's like wearing short pants in the summer tundra.

It's going to change a lot the internet we knew, unless somehow we managed to agree on a common high quality dataset repository.

Re: Gentoo bugzilla closed due AI bot scraper overload

#67
Enough with playing around the issue. The way out of this mess is not to protect with tech that works but with principles and laws.

This is not a tech problem. This is about what should or should not be legal.

Nor it’s a question of having time to implement solution X or Y.

If someone attacks you yes, you should have better security but you also need to have legal recourse, or it will never stop.

Ddos is already illegal.

I am not a lawyer so don’t ask me for exact resources, which vary by country anyway, but stop treating scrapers as an inescapable force of nature.

Re: Gentoo bugzilla closed due AI bot scraper overload

#68
post #56

I don't really understand why there is resistance to in-browser Anubis-style gating using crypto mining for public sites like this. You mine crypto and my free service remains cost-neutral. Yes, that transfers some of the cost to you, but it is marginal. You want to bang on my servers with the fury of a thousand madmen, then I'll scale up the servers and you can pay the marginal cost of your access. Or micropayments,…

Instead of punishing a criminal, you make everybody pay. This is not only unfair to legitimate users having to pay. It also needs to explain what actually happens when the Ddos succeeds (I know, you are talking about scrapers, but what’s the difference really?). Does that mean that the attacker just gets to shrug it off. Don’t take my statements as facts, I just want to outline a few reason I could come up with that…

If we could identify who is doing this, I bet that those people would be publicly burned at the stake by now.

Re: Gentoo bugzilla closed due AI bot scraper overload

#69

Earlier quoted context omitted.

How would I shut down a bot on Digital Ocean making over a million requests a day?

Send an email to the Digitalocean abuse email with the time ranges and origin IPs.

I did this with OVH, never got any response.

Re: Gentoo bugzilla closed due AI bot scraper overload

#70

Earlier quoted context omitted.

That would gate the internet to whole lower income countries

Aren't many of those countries basically being de-facto blocked by Cloudflare filtering out spam already anyways? Genuine question, I'm not up to date on how Cloudflare operates right now

The genuine answer to your question is no, they are not de-facto blocked, unless the website operator chooses to do so.
Post reply on HN