Earlier quoted context omitted.
Well, those companies have bots that identify themselves, and you can see what they're doing. Google especially have decades of experience of designing scrapers and seem to be able to the job of scraping the whole internet without causing problems in that time. So, it's still possible they are also running a bunch of poorly-coded scraper bots through residential proxies, but it seems a little non-obvious why they wou…
>So, it's still possible they are also running a bunch of poorly-coded scraper bots through residential proxies, but it seems a little non-obvious why they would be doing this. why "non obvious"? The easiest explanation is gathering training data sets, and is mostly caused by AI companies guarding their pile of essentially stolen IP (given how little they care about copyright) from eachother, without sharing any comp…
Gentoo bugzilla closed due AI bot scraper overload
101–110 of 114 posts
Re: Gentoo bugzilla closed due AI bot scraper overload
#102While we're dealing with the same issue at work, I sometimes still wonder exactly who these scrapers are. OpenAI, Google, Anthropic and others are normally fairly well behaved (Minus Anthropic attempting to hide behind a browser-for-hire company). Mostly you can get IP range and user-agents for the large players, while it problems mostly stem from bots pretending to be Chrome. Our largest offenders seems to be mostly…
You need to account for the fact that a lot of scraping is delegated by the big players to smaller players who can take loss of reputation and can use questionable methods (residential IPs). Some of the scrapers from these AI companies have been written very poorly from performance perspective. From what I understand the pressure has created even Google to be a lot more aggressive than what it was before. Not entirel…
> Residential proxies are everywhere, so why did proxy DDoS attacks mostly come from the U.S.? The answer is money. If you are committing fraud or circumventing content restrictions, a Russian IP address gets geo-blocked instantly. A fresh U.S. residential IP address (especially one behind carrier-grade NAT and harder to block individually) is “gold.” Customers pay up to $95 to lease a single U.S. residential IP address for 2 weeks (versus $0.30 for an Eastern European IP). Compare that to your own ARPU per subscriber and sit with it for a second. When an individual IP is worth more than the customer relationship behind it, you don’t have a technical problem. You have a market problem.[1]
1: https://www.nokia.com/blog/one-year-later-the-residential-pr...
Re: Gentoo bugzilla closed due AI bot scraper overload
#103Earlier quoted context omitted.
One of the defining characteristics of AI turned out to be utter and profound facelessness. The scraping, the data centers, the slop content, it's like it all comes from thin air. Nobody is talking about who is really doing all this or why, because it's always just coming from... somewhere else outside of our place.
And then it's pushed on the end users with the same amount of explanation and reason. You must adopt this now . The FOMO is insane.
do we have the same understanding of what fomo means?
Re: Gentoo bugzilla closed due AI bot scraper overload
#104Earlier quoted context omitted.
With micropayments, the server owner makes 1 million requests times 0.05 cents = ~$500 The status quo with those js PoW pages doesn't really benefit the server owner at all, it's wasted energy. I mean proof of work is always wasted energy, but I figure it's better to kill two birds with one stone. CoinHive was one example of this. (I think this is a correct link? https://github.com/cazala/coin-hive ). Although I thin…
$500 would amount to just 10k requests. One million requests is $50k
Re: Gentoo bugzilla closed due AI bot scraper overload
#105While we're dealing with the same issue at work, I sometimes still wonder exactly who these scrapers are. OpenAI, Google, Anthropic and others are normally fairly well behaved (Minus Anthropic attempting to hide behind a browser-for-hire company). Mostly you can get IP range and user-agents for the large players, while it problems mostly stem from bots pretending to be Chrome. Our largest offenders seems to be mostly…
You need to account for the fact that a lot of scraping is delegated by the big players to smaller players who can take loss of reputation and can use questionable methods (residential IPs). Some of the scrapers from these AI companies have been written very poorly from performance perspective. From what I understand the pressure has created even Google to be a lot more aggressive than what it was before. Not entirel…
Re: Gentoo bugzilla closed due AI bot scraper overload
#106While we're dealing with the same issue at work, I sometimes still wonder exactly who these scrapers are. OpenAI, Google, Anthropic and others are normally fairly well behaved (Minus Anthropic attempting to hide behind a browser-for-hire company). Mostly you can get IP range and user-agents for the large players, while it problems mostly stem from bots pretending to be Chrome. Our largest offenders seems to be mostly…
This has zero to do with AI (it's just a convenient scapegoat that fits the narrative). Absolutely everything to do with those pushing for centralised control of the Internet.
Re: Gentoo bugzilla closed due AI bot scraper overload
#107Re: Gentoo bugzilla closed due AI bot scraper overload
#108There are definitely patterns you can use against the scrapers. This maintainer just didn't have time for it, which is understandable. We direct scraper traffic to a bot-specific server using Cloudflare's load balancer, slowly analyzing traffic and adding conditions one at a time. No accidental scraper DDoS in a long time. Most scrapers are relatively honest in some way shape or form.
>We direct scraper traffic to a bot-specific server using Cloudflare's load balancer, slowly analyzing traffic and adding conditions one at a time. No accidental scraper DDoS in a long time. What's the point of this compared to letting cloudflare handle everything automagically? Presumably whatever heuristics they come up with are going to be better than you can, given limited time and budget?
To your point: I think you're simply mistaken that they offer a good scraper-detection solution, but I'm all ears.
Re: Gentoo bugzilla closed due AI bot scraper overload
#109Earlier quoted context omitted.
Something like Chaum's blind-signature based ecash, or GNU Taler would be ideal in terms of efficiency, but it's still centralized in distribution. Freenet currently uses something like this. Another option I was thinking of would be a pretty inflationary (or demurrage) cryptocurrency in which you have some sort of RandomX or other CPU-bound PoW. A web server could act as a mining pool and use mining shares interchan…
Lightning - a fast, instant, bitcoin layer 2 network - is perfectly sufficient for micropayments. Volatility is a no-issue in this case, as you can freely trade the 5 cents in realtime into other assets and minimize holding time of BTC. You will loose the spread, ofc.
1) it requires you to open a lightning channel. So you need some amount of initial investment (which usually means you need a credit card and an account on a bitcoin exchange), not just any computer which is the issue that using mining shares solves. The initial investment is my main grudge, as it is way too much friction for 402 Payment Required applications. Also as bitcoiners like sztorc note, this isn't feasible for most of the world's population due to bitcoin's small block size.
2) It's not private. People are going to be linking these things to their identities on crypto exchanges. So now feds can basicially track everything you do on the internet that way. LN is only private in the sense that not every transaction is broadcast to everyone else on the network, which is a very low bar. Chainalysis is possible with the right connections.
To a lesser extent, 3) Centralization in the lightning routing protocol which contributes to the effect of #2.
Re: Gentoo bugzilla closed due AI bot scraper overload
#110Enough with playing around the issue. The way out of this mess is not to protect with tech that works but with principles and laws. This is not a tech problem. This is about what should or should not be legal. Nor it’s a question of having time to implement solution X or Y. If someone attacks you yes, you should have better security but you also need to have legal recourse, or it will never stop. Ddos is already ille…
The question is, even if you identify who is the source of the traffic, are they even in a place where you can realistically sue them? The problematic traffic generally is not the bots that identify themselves, but the ones that are using residential proxies and try to be a non-fingerprintable as possible.
In whatever country they are.
I am not saying this is the magic bullet that will make it all go away. Some will, some will not.
Enforcement’s aim is not to completely eradicate crime: make it anti-economical for the majority and it’s a good deal enough.
Go to Berlin and torrent a movie: you will be fined, for a few hundred euros. Can you VPN and get away with it? I suppose so, but your average Joe is just not going to do it: why pay 10 euro per month on a VPN and risk? Just rent it.
DDOS is also, in general, a criminal charge. What would happen to prices once you get a few inducements? If it’s a crime and enforced, what are the roles of App Stores that let users download residential proxy? Are they also committing a crime? Residential proxy can be used for legitimate purposes? I don’t think that line of defence worked well for Pirate Bay, no?
And yes pirate bay is still around, some scraping will always be there.