Live data from Hacker News

An update on residential proxies and the scraper situation

lwn.net

391–400 of 422 posts

Re: An update on residential proxies and the scraper situation

#391

Earlier quoted context omitted.

HN is surprisingly very very guilty of a whole lot of anti-user patterns and behaviour that other companies get regularly lamented. Poor accessibility, bad mobile support, no options to delete content beyond a narrow window.

Poor accessibility Good.

Only for ghouls.

Re: An update on residential proxies and the scraper situation

#392

Earlier quoted context omitted.

That's a completely different question, your claim was about parallelism.

Ah, I see what you're getting at. Yes. You can think of any given computer aa having a fixed amount of compute budget for these types of acceleration-resistant hashing algorithms. Let's say the scraper can perform 10,000 hashing operations per second total on their machine, and needs an average of 1,000 hashing operations to solve the PoW. It's a minor detail, but note that these PoW challenges non-deterministically…

So, in your example, each request needs 0.1 core-seconds of PoW. Is that right? Or, since you're saying 10000 hashes across the entire machine and 1000 needed, then on a typical 96-core server you need 9.6 core-seconds of PoW? The former means you pay about $0.0000005 per request at standard cloud rates; the latter means you pay about $0.00005 per request _and_ your site is totally unusable by legitimate clients. Both are _easily_ worth it for someone backed by VC billions and hungry for data. The network and storage fees are likely to be more significant than that already, not to mention actually training a model; they don't balk at downloading terabytes of crap already.

Note that none of this assumes any sort of acceleration from parallelization (be it through GPUs or reusing work across servers) or precomputation relative to what a normal client does. Compute is just really cheap in dollars compared to the cost of having a user wait, and these companies _also_ have a lot of appetite for spending dollars compared to that of a normal user. As others have pointed out, the only reason why Anubis works (sort-of; not for everyone) right now is that it is uncommon enough, essentially “proof that you bothered to have your crawler run JavaScript at all”. It's a confusion measure.

Proof of work does not work.

Re: An update on residential proxies and the scraper situation

#393

Earlier quoted context omitted.

I think they also have to operate in countries that don't mind shady things like this.

Residential proxies aren't illegal in any country, within reason.

They're shady though. Open to civil litigation, certainly. No big player would touch this operating in US.

Re: An update on residential proxies and the scraper situation

#394

Earlier quoted context omitted.

Ah, I see what you're getting at. Yes. You can think of any given computer aa having a fixed amount of compute budget for these types of acceleration-resistant hashing algorithms. Let's say the scraper can perform 10,000 hashing operations per second total on their machine, and needs an average of 1,000 hashing operations to solve the PoW. It's a minor detail, but note that these PoW challenges non-deterministically…

So, in your example, each request needs 0.1 core-seconds of PoW. Is that right? Or, since you're saying 10000 hashes across the entire machine and 1000 needed, then on a typical 96-core server you need 9.6 core-seconds of PoW? The former means you pay about $0.0000005 per request at standard cloud rates; the latter means you pay about $0.00005 per request _and_ your site is totally unusable by legitimate clients. Bot…

You raise some good points here, but PoW is meant to be one tool in the toolbox, not the only line of defense. You can still maintain blocklists of known scrapers (or better yet, have your PoW system be aware of them and silently adjust the difficulty to an impossible level, such that the scraper gets stuck trying to solve your PoW challenge until it hits a timeout configured by the scraper, if they were wise enough to configure one). It's also courteous to not only build and maintain your own blocklists, but to share them with e.g. vtotal and spamhaus, to help protect others.

Similarly, you have tarpits, which generate infinite mazes of garbage data, or even deliberately poisoning training data (should the scrapers be training LLMs) though this don't entirely eliminate the deleterious effects of scraping on the host's web server (more info: https://arstechnica.com/tech-policy/2025/01/ai-haters-build-...).

If the premise was evaluating whether or not PoW would be a magic silver bullet that stops scrapers all by itself, then you are correct, it does not stop all scrapers. Scraping and anti-scraping is fundamentally a constantly evolving cat and mouse game that demands adaptability and punishes complacency from all participants trying not to lose.

Re: An update on residential proxies and the scraper situation

#395

The article at the end talks about how is very easy for arbitrary apps from app stores can install a residential proxy on your phone. 10 years ago, apps had to explicitly state if they needed network access. And then the powers that be decided that really all apps need network access no matter what. And both ios and android make it hard to deny apps network access. But really, this finally explains the hordes of real…

Google should be able to detect this and ban those apps from the Play Store. They have the incentive too.

Re: An update on residential proxies and the scraper situation

#396
post #10

Residential Proxies are the most emblematic technology of our era- a group of people looked at something that used to be considered a crime (botnets) and realized that if they just did it openly, no one would ever punish them.

This is a super dishonest characterization. Running software on a bunch of machines, even machines in other peoples' homes has never been a crime. Folding@home isn't a crime (obviously). It's controlling those machines without consent via malware that is criminal. And if it is open and consensual in exchange for something a person wants, it is unreasonable to compare it to botnets.

Software that implements residential proxy networks is always covert and never consensual.

It doesn't matter if page 42 of the Terms and Condition "clearly states it." This does not constitute informed consent.

Re: An update on residential proxies and the scraper situation

#397

Earlier quoted context omitted.

Residential proxies aren't illegal in any country, within reason.

They're shady though. Open to civil litigation, certainly. No big player would touch this operating in US.

Good thing we have small players who are still willing to innovate.

Re: An update on residential proxies and the scraper situation

#398

Earlier quoted context omitted.

Noone complained we can't discriminate DC IPs, though to be fair some (imo bad) operators did just that. This is not even about preventing bots, which has perfectly legitimate usecases (eg. Internet Archive). This is about filtering out bad bots/actors who have no respect for your resources and will drain all of it causing bad experience for everyone. But because they know they don't respect robots.txt or even simple…

You could always try various counterattacks - returning a repeating compressed stream that decompresses to several terabytes, or returning one byte per second, or returning an endless list of hyperlinks to a honeypot, if the date is older than 200 years.

Sure, a lot of countermeasures is possible. I will definitely consider it next time, but for this calendar which had sitted unused for years on this instance, i decided to just turn off the plugin. A random wordpress plugin is never only a DoS avenue but also a bunch of potential vulns.

Thanks for the suggestions!

Re: An update on residential proxies and the scraper situation

#399
post #345

Earlier quoted context omitted.

The success of the ongoing Anubis rollout proves the opposite. People are used to slowly-loading websites - the rise of garbage SPAs has seen to that. Staring at a spinner for a second every once in a while is not an issue for genuine users. On the other hand, the additional CPU usage rises the compute cost of scraping by several orders of magnitude. If you don't have to scrape this specific website, you'd be stupid…

Scrapers generally aren't looking for random websites to scrape. They have a specific URL in mind. Only if the goal was DDoS would they not care which URL was accessed.

They also don't care which URL was accessed when the goal is scraping as much text as possible for AI training.

Re: An update on residential proxies and the scraper situation

#400
post #325
post #287

Earlier quoted context omitted.

Why is that good?

It's relatively difficult to enter and edit any significant amount of text on a phone, so phone-based discussions tend to be shallow.

That’s a terrible justification, and I would question if it’s even true. Is there anything to support that conclusion or is it just a guess you’ve made?

Because I don’t buy it. There is no time limit for discussions. One can take as much time as they like to enter text on their phone. Also if they’re motivated by the discussion they may take the time and make the effort to write thoughtful replies.

Finally, there’s no reason that users visiting on desktop computers won’t make equally shallow remarks. Shallow online discussions have existed long before smartphones existed.

Sent from my iPhone.

Post reply on HN