Live data from Hacker News

An update on residential proxies and the scraper situation

lwn.net

381–390 of 422 posts

Re: An update on residential proxies and the scraper situation

#381
post #348

As this article points out, it's tremendously unclear who is using residential proxies . The big AI models claim they're not using them. I'm not inclined to "just believe them", but no incriminating evidence has leaked, and—as pointed out in the article—many of the bots that are running on these residential proxy botnets are coded in incredibly stupid and inefficient ways. How confident are people who research this s…

That's the only theory we have that doesn't sound like a conspiracy theory. The only other credible ideas are that someone's doing some kind of dataset arbitrage by scraping the fuck out of everything and selling companies data that is technically new (by means of the scrape date being newer).

Right, and as I understand it the timing also lines up: these ill-behaved scraper-bot-nets exploded along with GenAI in the last 3-4 years.

Still, it seems to me like something doesn't add up. Running these botnets is perhaps cheap but it isn't free, and dumping all their data into LLM training is truly expensive: would so many of these bots be so blitheringly inefficient in their scraping patterns if all of the results were getting fed into LLM training?

Have there been no leaks or whistleblowers from the "semi-legit" RP brokers?

Re: An update on residential proxies and the scraper situation

#382

Earlier quoted context omitted.

Regardless of the precise logic it's no excuse for the policy. Simply hand out a sufficiently difficult PoW to prevent widespread abuse. I'm quite certain it isn't a generic "datacenter" list though because a given VPN exit that was working will suddenly stop. Meanwhile I have a valid cookie yet that is disregarded.

Wikipedia maintains a quite exhaustive list of them to prevent spam and they also block vpn IPs as soon as they're found.

That is an entirely different matter. They only block (AFAIK) editing which is a scenario where PoW would not be expected to solve the problem. What I was talking about here is IP banning as a (boneheaded) mitigation against high volume scrapers.

Re: An update on residential proxies and the scraper situation

#383

Earlier quoted context omitted.

Cloudflare often just straight up blocks me or makes me do a captcha. IMO those are both much worse than Anubis

Sounds like it's because your IP is on a data Center IP list

That's lame of them, I'm on a residential plan. Maybe because I host my personal website from home?

Re: An update on residential proxies and the scraper situation

#384
post #319
post #299

Earlier quoted context omitted.

We get thousands of emails - far too many to even read them all, let alone respond the way we would like to. There are other reasons why we might not respond, but overload is the main one. Edit: I only see one email in the archive related to your account. It was from Sept 2024 and we responded to it. Are you talking about a different account?

Oh, okay I didn't know you weren't getting around to all incoming emails, that makes sense then. Previously the responses were fast and reliable and I thought I was being dropped by a silent bot/spam detection algorithm (as is becoming frustratingly common). It's indeed from two different email addresses from the one on this account. They're not that important so never mind, thanks for checking and for the reply!

If you give us something to look up I can try to track them down for you. Up to you of course.

Re: An update on residential proxies and the scraper situation

#385
post #69

Earlier quoted context omitted.

It's really bad. I found myself identifying with everything Jonathan wrote in the OP - so much so that I thought of asking to compare notes on mitigation measures.

I came up with the idea and helped build Grub, the distributed crawler. Looksmart bought it, ran it for a time, then sold it to Wikimedia. I reclaimed the name recently (abandoned mark) and have a new crawler now that is agentic. I use it for my own research runs, and it's not my main focus at this point, nor am I trying to get it attention. A lot of LLMs and coding agents can easily fetch content if it is needed and…

[deleted]

Re: An update on residential proxies and the scraper situation

#386
I don't run one of these sites that has these issues so I'm really not aware of this problem. How can it be that sites are getting overwhelmed with scrapers that are just looking for training data? You only need to scrape it once to train a model, so shouldn't there be less traffic from this than there is from search engines?

On the other hand if the article is wrong and the traffic is coming from other ai uses (like an agent visiting pages on behalf of a user) then that would make sense.

Re: An update on residential proxies and the scraper situation

#388

Earlier quoted context omitted.

HN is surprisingly very very guilty of a whole lot of anti-user patterns and behaviour that other companies get regularly lamented. Poor accessibility, bad mobile support, no options to delete content beyond a narrow window.

no options to delete content beyond a narrow window. Good.

hurfdurf says it is okay to keep material online forever. It is easy to say that when you publish as hurfdurf.

Re: An update on residential proxies and the scraper situation

#389

The article at the end talks about how is very easy for arbitrary apps from app stores can install a residential proxy on your phone. 10 years ago, apps had to explicitly state if they needed network access. And then the powers that be decided that really all apps need network access no matter what. And both ios and android make it hard to deny apps network access. But really, this finally explains the hordes of real…

> 10 years ago, apps had to explicitly state if they needed network access. And then the powers that be decided that really all apps need network access no matter what. Why does network access need to be a binary, all or nothing? When you install an app, the app should request permissions to specific DNS names, i.e. pointing to the servers that the app's authors operate. If I install Todoist, the app should only ask…

[deleted]

Re: An update on residential proxies and the scraper situation

#390
post #271

Earlier quoted context omitted.

HN is surprisingly very very guilty of a whole lot of anti-user patterns and behaviour that other companies get regularly lamented. Poor accessibility, bad mobile support, no options to delete content beyond a narrow window.

bad mobile support Good.

If Bad Mobile support is good why does your own website handles mobile differently?
Post reply on HN