Live data from Hacker News

Disrupting the largest residential proxy network

cloud.google.com

161–170 of 230 posts

Re: Disrupting the largest residential proxy network

#161

Earlier quoted context omitted.

Am I the only one cynically thinking that "Russia, Iran, DPRK, PRC, etc" is the "But think of the chiiildren!!!" excuse for doing this? And when Google say "IPIDEA’s proxy infrastructure is a little-known component of the digital ecosystem leveraged by a wide array of bad actors." What they really mean is " ... leveraged by actors indiscriminately scraping the web and ignoring copyright - that are not us." I can't he…

No, what they're saying is what they said, what you're implying reveals a strange bias. Web scraping through residential proxies? Please think through your thoughts more. There's much more effective and efficient ways to do so. Multiple bad actors, like ransomware affiliates, have been caught using residential proxy networks. But by all means, don't let facts and cyber threat intelligence get in the way.

>let facts and cyber threat intelligence get in the way

Appeal to authority by way of invoking the megacorp-branded "threat intelligence" capability (targeted PR exercise).

Re: Disrupting the largest residential proxy network

#163

Residential proxies are the only way to crawl and scrape. It's ironic for this article to come from the biggest scraping company that ever existed! If you crawl at 1Hz per crawled IP, no reasonable server would suffer from this. It's the few bad apples (impatient people who don't rate limit) who ruin the internet for both users and hosters alike. And then there's Google.

I'd still like the ability to just block a crawler by its IP range, but these days nope.

1 Hz is 86400 hits per day, or 600k hits per week. That's just one crawler.

Just checked my access log... 958k hits in a week from 622k unique addresses.

95% is fetching random links from u-boot repository that I host, which is completely random. I blocked all of the GCP/AWS/Alibaba and of course Azure cloud IP ranges.

It's almost all now just comming of a "residential" and "mobile" IP address space from completely random places all around the world. I'm pretty sure my u-boot fork is not that popular. :-D

Every request is a new IP address, and available IP space of the crawler(s) is millions of addresses.

I don't host a popular repo. I host a bot attraction.

Re: Disrupting the largest residential proxy network

#164

We need more residential proxies, not less. I've had enough of companies saying "you're connecting from an AWS IP address, therefore you aren't allowed in, or must buy enterprise licensing". Reddit is an example which totally blocks all data to non-residential IP's. I want exactly the same content visible no matter who you are or where you are connecting from, and a robust network of residential proxies is a stepping…

> I want exactly the same content visible no matter who you are or where you are connecting from The reason those IP addresses get blocked is not because of "who" is connecting, but "what" Traffic from datacenter address ranges to sites like Reddit is almost entirely bots and scrapers. They can put a tremendous load on your site because many will try to run their queries as fast as they can with as many IPs as they c…

Okay. So what does ten million requests cost, then? Like... a dollar? Is it a dollar? Is it two dollars if they splurge?

Because if the deterrent here is a line item so small it shows up as 'miscellaneous vibes' on a balance sheet, that's not a barrier. That's a tip jar.

Re: Disrupting the largest residential proxy network

#165

Earlier quoted context omitted.

Am I the only one cynically thinking that "Russia, Iran, DPRK, PRC, etc" is the "But think of the chiiildren!!!" excuse for doing this? And when Google say "IPIDEA’s proxy infrastructure is a little-known component of the digital ecosystem leveraged by a wide array of bad actors." What they really mean is " ... leveraged by actors indiscriminately scraping the web and ignoring copyright - that are not us." I can't he…

>I can't help but feel this is just Google trying to pull the ladder up behind then and make it more difficult for other companies to collect training data. I can very easily see this as being Google's reasoning for these actions, but let's not pretend that clandestine residential proxies aren't used for nefarious things. The vast majority of social media networks will ban - or more generally and insiously - shadow b…

> The vast majority of social media networks will ban - or more generally and insiously - shadow ban accounts/IPs that use known proxy IPs. This means that they are gating access to their platforms behind residential IPs (on top of their other various blackboxes and heuristics like fingerprinting)

Social media will ban proxy IPs, yet gleefully force you to provide your ID if you happen to connect from the wrong patch of land. I find it difficult not to support any and all attempts to bypass such measures.

The fact is that there's now a perfectly legitimate use for residential proxies, and the demand is just going to keep growing as more websites decide to "protect their content", and more governments decide to pass tyrannical laws that force people to mask their IPs. And with demand, comes supply, so don't expect them to go away any time soon.

This really just sounds like a rehash of the argument against encryption. "Bad people use it, so it should go away" - never mind that there are completely legitimate uses for it. Never mind that using a residential proxy might be the only way to get any privacy at all in a future where everyone blocks VPNs and Tor, a future where you may not even be able to post online without an ID depending you where you live, a future which we're swiftly approaching.

It's already here, in fact. Imgur blocks UK users, but it also blocks VPNs and Tor. The only way somebody living in the UK can access Imgur is through a residential proxy.

Re: Disrupting the largest residential proxy network

#166

Residential proxies are the only way to crawl and scrape. It's ironic for this article to come from the biggest scraping company that ever existed! If you crawl at 1Hz per crawled IP, no reasonable server would suffer from this. It's the few bad apples (impatient people who don't rate limit) who ruin the internet for both users and hosters alike. And then there's Google.

First of: Google has not once crashed one of our sites with GoogleBot. They have never tried to by-pass our caching and they are open and honest about their IP ranges, allowing us to rate-limit if needed.

The residential proxies are not needed, if you behave. My take is that you want to scrape stuff that site owners do not want to give you and you don't want to be told no or perhaps pay a license. That is the only case where I can see you needing a residential proxies.

Re: Disrupting the largest residential proxy network

#167
post #131
post #122

Earlier quoted context omitted.

Users are OK with acting as proxies because they don't understand all the shady stuff their proxy is being used for. Also consumer ISPs generally ban this.

But then would you make the same arguments for running a tor node (presumably, you don't know what shady stuff is there, but you know there's shady stuff)?

That's totally something you should consider, even if you decide for running the tor node anyway in the end.

Re: Disrupting the largest residential proxy network

#168
post #25

so that only google and anthropic are allowed to scrape the web. No one else may have workarounds

Anyone could scrape the net, then modern scrapes came along with their shitty code and absolutely no respect. The reason why so many of us block or throttle scrapers is because they miss behave. They don't back off, they try to by-pass caches and if they crash a site they don't adjust, they will just pound it the ground again when it's back. We managed to talk to one large AI company would didn't really want to fix anything, but told us that they'd be fine with us just rate limiting them, as if we somehow owed them anything. They just get a stupid low rps now, even if we'd let them go faster, if they'd just fix they bot.

Some sites don't want you scraping, but it's their content, their rules. We don't really care, but we have to due to the number and quality of the bots we're seeing. This is in my mind a 100% self-imposed problem from the scrapers.

Re: Disrupting the largest residential proxy network

#170
post #131
post #122

Earlier quoted context omitted.

Users are OK with acting as proxies because they don't understand all the shady stuff their proxy is being used for. Also consumer ISPs generally ban this.

But then would you make the same arguments for running a tor node (presumably, you don't know what shady stuff is there, but you know there's shady stuff)?

Running a tor node is pretty stupid from a liability perspective, but at least you have more deniability and you are making an informed choice.

These residential proxies are pretty much universally shady. I doubt most of the users understand what they are consenting to.

Post reply on HN