> ...we have tried to minimize the impact on real readers as much as possible. We have not gone with tools like Anubis, partly because it causes annoying delays for those trying to get to the site, but also partly because it seems inevitable that the scrapers will eventually find their way around it. Indeed, there are some indications that is already happening. A proof-of-work requirement is not a huge obstacle when…
PoW barely affects the "residential proxies" aka. malware botfarms. The IPs are free for them and siphoning additional system resources for PoW doesn't matter at all for them. PoW only affects the large centralised scraping by the AI providers, which are not operating behind "residential proxies".
An update on residential proxies and the scraper situation
211–220 of 422 posts
Re: An update on residential proxies and the scraper situation
#212I feel like the solution is a better common crawl. As nice as it would be to block the frontier AI labs from getting access to information, we should reset the baseline of information accessibility so there's less marginal advantage on these labs. I worry a lot of the anti scraping rhetoric will just injure the open web and put somebody like cloudflare in charge.
What really confuses me is ... people always say, it's because companies are gathering data for AI training. Then why would they need to scrape the same page thousands of times per day? Edit: the article says millions of times per hour? (!?) The article is also astonished by this, and speculates it might be some kind of underground AI labs but... millions of them? Or does it only take one with too much money and a ba…
Re: An update on residential proxies and the scraper situation
#213Earlier quoted context omitted.
As someone dealing with similar on a large site, I'd love to see a private community to discuss some of these issues.
Disclaimer: this works for my very small number of personal services that I run. I have no idea how this would (or probably wouldn't) scale at all. Also, the methodology I describe below is based on what I'm able to do technically, which is pretty much limited to bash scripting. On my external-most device I have a firewall that logs addresses that attempt to connect to ports behind which there are no services, and th…
Re: An update on residential proxies and the scraper situation
#214In theory, proof of work that is used to mine a cryptocurrency could be a solution. Bitcoin and others are already secured via massive pow computations. If we could shift that into browsers, no additional energy would be used and we could solve an issue that has been unsolved for too long: How to pay websites that provide useful information other than with ads. The question is which resources typical consumer hardwar…
Proof of work, even "custom", where the user does not need a particular interaction with the page, does not work. The scrapers are running headless Chrome and solving the work. They do not care, they do not pay the bill, the compromised system's owner pays the bill. I have such system for the registration form on one of my website to prevent the double validation of emails to be used to spam emails of victims. The Po…
Re: An update on residential proxies and the scraper situation
#215In theory, proof of work that is used to mine a cryptocurrency could be a solution. Bitcoin and others are already secured via massive pow computations. If we could shift that into browsers, no additional energy would be used and we could solve an issue that has been unsolved for too long: How to pay websites that provide useful information other than with ads. The question is which resources typical consumer hardwar…
Proof of work captchas are widely deployed, especially on more niche sites. Kiwiflare is one (used for a harassment forum)
Re: An update on residential proxies and the scraper situation
#216Earlier quoted context omitted.
Yes I've seen it. ClaudeBot will gleefully announce itself when it hammers my niche website a million times a day.
At least those bots are easy to block though. I run a niche stats website for an esport and I have no idea why there's loads of residential trawlers/botnets with 10k+ IPs trying to get that data - most of what they scrape is directly available from Valve's APIs.
Re: An update on residential proxies and the scraper situation
#217Residential Proxies are the most emblematic technology of our era- a group of people looked at something that used to be considered a crime (botnets) and realized that if they just did it openly, no one would ever punish them.
Well, one party gives away free walls if you agree to fill your castle with surveillance cameras you don't control.
Re: An update on residential proxies and the scraper situation
#218Earlier quoted context omitted.
by consent they mean a dialog/EULA with careful wording was display and the user clicked ok.
And even if you spent a minute explaining the proposition to a user off the street, it still wouldn't be fair unless you laid out the drawbacks. Which leads me to a question. There must be countless individuals all over the world who suddenly can't log into their Gmail or create any new accounts because a fraudster sent spam from their IP. I wonder: has anyone has tried to quantify that problem?
Re: An update on residential proxies and the scraper situation
#219Earlier quoted context omitted.
Thank god for residential proxies. Highly unethical but the way the internet is going they're the last anti-hero of a somewhat open internet
By providing a way for corporate AI scrapers to operate with impunity and force the last few independently-run websites to move to the cloud?