Live data from Hacker News

Disrupting the largest residential proxy network

cloud.google.com

221–230 of 230 posts

Re: Disrupting the largest residential proxy network

#221

Earlier quoted context omitted.

> The vast majority of social media networks will ban - or more generally and insiously - shadow ban accounts/IPs that use known proxy IPs. This means that they are gating access to their platforms behind residential IPs (on top of their other various blackboxes and heuristics like fingerprinting) Social media will ban proxy IPs, yet gleefully force you to provide your ID if you happen to connect from the wrong patch…

> The only way somebody living in the UK can access Imgur is through a residential proxy. And very little of value was lost. > This really just sounds like a rehash of the argument against encryption. "Bad people use it, so it should go away" - never mind that there are completely legitimate uses for it. Except that almost everything that uses encryption has some legitimate use. There are pretty much no legitimate us…

It's another type of proxy. Legitimate uses are the same as for other types of proxies.

Re: Disrupting the largest residential proxy network

#222
post #131
post #122

Earlier quoted context omitted.

Users are OK with acting as proxies because they don't understand all the shady stuff their proxy is being used for. Also consumer ISPs generally ban this.

But then would you make the same arguments for running a tor node (presumably, you don't know what shady stuff is there, but you know there's shady stuff)?

Personally I consider Tor less shady than these residential proxy networks because Tor has some normal users but yes, the considerations are similar. (I ran one of the earliest Tor exit nodes.)

Re: Disrupting the largest residential proxy network

#223
post #163

Residential proxies are the only way to crawl and scrape. It's ironic for this article to come from the biggest scraping company that ever existed! If you crawl at 1Hz per crawled IP, no reasonable server would suffer from this. It's the few bad apples (impatient people who don't rate limit) who ruin the internet for both users and hosters alike. And then there's Google.

I'd still like the ability to just block a crawler by its IP range, but these days nope. 1 Hz is 86400 hits per day, or 600k hits per week. That's just one crawler. Just checked my access log... 958k hits in a week from 622k unique addresses. 95% is fetching random links from u-boot repository that I host, which is completely random. I blocked all of the GCP/AWS/Alibaba and of course Azure cloud IP ranges. It's almos…

In addition to a rate limit, a page limit per IP is needed; this is specifically for things like source code repos (with massive commit histories), mailing archives, etc.

A whitelist would be needed for sites where getting all the pages make sense. And probably in addition to the 1Hz, an additional limit of 1k/day would be needed.

I can see now why Google has not much solid competition (Yandex/Baidu arguably don't compete due to network segmentation).

Scraping reliably is hard, and the chance of kicking Google off their throne may be even further reduced due to AI crawler abuse.

PS 958k hits is a lot! Even if your pages were a tiny 7.8k each (HN front page minus assets), that would be about 7G of data (about 4.6 Bee Movies in 720p h256).

Re: Disrupting the largest residential proxy network

#224

Earlier quoted context omitted.

No, the question is not just disclosure. People have their bandwidth stolen, and sometimes internet access revoked due to this kind of fraud and misuse - disclosure wouldn’t solve that

Also, as a website owner, these residential proxies are a real pain. Tons and tons of abusive traffic, including people trying to exploit vulnerabilities and patently broken crawlers that send insane numbers of requests, and no real way to block it. It's just nasty stuff. Intent matters, and if you're selling a service that's used only by the bad guys, you're a bad guy too. This is not some dual-use, maybe-we-should-…

I run a really small forum and I've been absolutely inundated with a bunch of junk traffic. I had to tighten my Cloudflare WAF rules a whole bunch, and start issuing browser challenges way more aggressively.

Excluding known "good" crawlers, well over 99% of the traffic trying to hit the site has been attempting to maliciously scrape. Most of this traffic looks genuine, but has random genuine-looking user agents and comes from random residential proxies in various countries, usually the US.

For the traffic that does make it all the way to a browser challenge, the success rate is a measly 0.48%. Put another way, over 50% of traffic is already blocked by that point, and of the under 50% that makes it to a browser challenge, more than 99.5% fails that challenge.

It's been virtually no disruption to users either, since I configured successful challenges to be remembered for a long period of time. The legitimate traffic is a gentle trickle, while the WAF is holding back garbage traffic that's orders of magnitude above and beyond normal levels. The scale of it is truly insane.

Re: Disrupting the largest residential proxy network

#225

Earlier quoted context omitted.

what industry is that? Every industry is on the cloud.

No, no they really aren't, but I was thinking the "scraping industry" in the sense that that's a thing. Getting hosting in smaller datacenters is simple enough, but you may need to manage your own hardware, or VMs. Many will help you get your own IP ranges and ASN, that's going to go a long way, if you don't want to get bundled in with the bad bots. This differs obviously, but having an ASN in our case means that we…

Thank you for speaking some sense. As a site operator that's been inundated with junk traffic over the past ~month where well in excess of 99% of it has to be blocked, the scrapers have brought this upon themselves.

I actually do let quite a few known, "good" scrapers scrape my stuff. They identify themselves, they make it clear what they do, and they respect conventions like robots.txt.

These residential proxies have been abused by scrapers that use random legit-looking user agents and absolutely hammer websites. What is it with these scrapers just not understanding consent? It's gross.

Re: Disrupting the largest residential proxy network

#226

Earlier quoted context omitted.

No, no they really aren't, but I was thinking the "scraping industry" in the sense that that's a thing. Getting hosting in smaller datacenters is simple enough, but you may need to manage your own hardware, or VMs. Many will help you get your own IP ranges and ASN, that's going to go a long way, if you don't want to get bundled in with the bad bots. This differs obviously, but having an ASN in our case means that we…

Scraping isn’t an industry. There are legitimate and illegitimate scraping pursuits. There are lots of healthy / productive businesses in the cloud and lots of scumbags, just like any enterprise. I still have no idea about your point, by the way.

If you have a "legitimate scraping pursuit", identify yourself appropriately that way. I'm happy to let most well-behaved scrapers access my content.

Hiding behind a residential proxy and using random user agents? Gross. Learn what consent is.

Re: Disrupting the largest residential proxy network

#227

Earlier quoted context omitted.

Does malicious mean interfering with Google's business model, or does it include intrusive advertising?

Malicious here means "most people who aren't trying to argue semantics or otherwise be smartasses about it would consider it malware". That's why the example I gave is a semi-popular software the allows watching YouTube without ads without a premium subscription, i.e. at least in the case I observed, I don't believe this was weaponized against apps that interfere with their business model. As for "intrusive advertisi…

I'm not being a smartass. Intrusive ads are malware. Adware used to be a category in virus scanners, then stopped when virus scanners wanted to run ads themselves.

Proxying traffic is not malware, since it doesn't affect me in any way.

Re: Disrupting the largest residential proxy network

#228
post #96

Earlier quoted context omitted.

Getting rid of malware is good. A private for-profit company exercising its power over the Internet, not so much. We should have appropriate organizations for this.

The proxies is the reason why you get spam in your Google search result, spam in your Play store (by means of fake good reviews), basically spam in anything user generated. It directly affects Google and you, I don’t see why they should not do this.

They are not the reason. They may be one mechanism used for this.

Re: Disrupting the largest residential proxy network

#229

Earlier quoted context omitted.

Scraping isn’t an industry. There are legitimate and illegitimate scraping pursuits. There are lots of healthy / productive businesses in the cloud and lots of scumbags, just like any enterprise. I still have no idea about your point, by the way.

If you have a "legitimate scraping pursuit", identify yourself appropriately that way. I'm happy to let most well-behaved scrapers access my content. Hiding behind a residential proxy and using random user agents? Gross. Learn what consent is.

try scraping any of the major players e.g. Amazon without residential proxy it won't work. I appreciate that you are offering to abide by crawling etiquette (e.g. robots.txt) but no major app supports that any more.

You're thinking about the case of big AI companies crawling your blog. I'm talking about a small startup trying to do traditional indexing and needing to run from residential proxy to make it work.

Re: Disrupting the largest residential proxy network

#230

Earlier quoted context omitted.

> Malware in random apps running on your device without your knowledge is bad. And ones that have all the indicators of compromise of Russia, Iran, DPRK, PRC, etc

Am I the only one cynically thinking that "Russia, Iran, DPRK, PRC, etc" is the "But think of the chiiildren!!!" excuse for doing this? And when Google say "IPIDEA’s proxy infrastructure is a little-known component of the digital ecosystem leveraged by a wide array of bad actors." What they really mean is " ... leveraged by actors indiscriminately scraping the web and ignoring copyright - that are not us." I can't he…

If they said "could" then I would agree but they said it did happen. those actors DID do it, not could. So it's not a think of the children excuse. Unless they are outright lying but I doubt the security team came up with a business type excuse
Post reply on HN