Earlier quoted context omitted.
> The vast majority of social media networks will ban - or more generally and insiously - shadow ban accounts/IPs that use known proxy IPs. This means that they are gating access to their platforms behind residential IPs (on top of their other various blackboxes and heuristics like fingerprinting) Social media will ban proxy IPs, yet gleefully force you to provide your ID if you happen to connect from the wrong patch…
> The only way somebody living in the UK can access Imgur is through a residential proxy. And very little of value was lost. > This really just sounds like a rehash of the argument against encryption. "Bad people use it, so it should go away" - never mind that there are completely legitimate uses for it. Except that almost everything that uses encryption has some legitimate use. There are pretty much no legitimate us…
Disrupting the largest residential proxy network
221–230 of 230 posts
Re: Disrupting the largest residential proxy network
#222Earlier quoted context omitted.
Users are OK with acting as proxies because they don't understand all the shady stuff their proxy is being used for. Also consumer ISPs generally ban this.
But then would you make the same arguments for running a tor node (presumably, you don't know what shady stuff is there, but you know there's shady stuff)?
Re: Disrupting the largest residential proxy network
#223Residential proxies are the only way to crawl and scrape. It's ironic for this article to come from the biggest scraping company that ever existed! If you crawl at 1Hz per crawled IP, no reasonable server would suffer from this. It's the few bad apples (impatient people who don't rate limit) who ruin the internet for both users and hosters alike. And then there's Google.
I'd still like the ability to just block a crawler by its IP range, but these days nope. 1 Hz is 86400 hits per day, or 600k hits per week. That's just one crawler. Just checked my access log... 958k hits in a week from 622k unique addresses. 95% is fetching random links from u-boot repository that I host, which is completely random. I blocked all of the GCP/AWS/Alibaba and of course Azure cloud IP ranges. It's almos…
A whitelist would be needed for sites where getting all the pages make sense. And probably in addition to the 1Hz, an additional limit of 1k/day would be needed.
I can see now why Google has not much solid competition (Yandex/Baidu arguably don't compete due to network segmentation).
Scraping reliably is hard, and the chance of kicking Google off their throne may be even further reduced due to AI crawler abuse.
PS 958k hits is a lot! Even if your pages were a tiny 7.8k each (HN front page minus assets), that would be about 7G of data (about 4.6 Bee Movies in 720p h256).
Re: Disrupting the largest residential proxy network
#224Earlier quoted context omitted.
No, the question is not just disclosure. People have their bandwidth stolen, and sometimes internet access revoked due to this kind of fraud and misuse - disclosure wouldn’t solve that
Also, as a website owner, these residential proxies are a real pain. Tons and tons of abusive traffic, including people trying to exploit vulnerabilities and patently broken crawlers that send insane numbers of requests, and no real way to block it. It's just nasty stuff. Intent matters, and if you're selling a service that's used only by the bad guys, you're a bad guy too. This is not some dual-use, maybe-we-should-…
Excluding known "good" crawlers, well over 99% of the traffic trying to hit the site has been attempting to maliciously scrape. Most of this traffic looks genuine, but has random genuine-looking user agents and comes from random residential proxies in various countries, usually the US.
For the traffic that does make it all the way to a browser challenge, the success rate is a measly 0.48%. Put another way, over 50% of traffic is already blocked by that point, and of the under 50% that makes it to a browser challenge, more than 99.5% fails that challenge.
It's been virtually no disruption to users either, since I configured successful challenges to be remembered for a long period of time. The legitimate traffic is a gentle trickle, while the WAF is holding back garbage traffic that's orders of magnitude above and beyond normal levels. The scale of it is truly insane.
Re: Disrupting the largest residential proxy network
#225Earlier quoted context omitted.
what industry is that? Every industry is on the cloud.
No, no they really aren't, but I was thinking the "scraping industry" in the sense that that's a thing. Getting hosting in smaller datacenters is simple enough, but you may need to manage your own hardware, or VMs. Many will help you get your own IP ranges and ASN, that's going to go a long way, if you don't want to get bundled in with the bad bots. This differs obviously, but having an ASN in our case means that we…
I actually do let quite a few known, "good" scrapers scrape my stuff. They identify themselves, they make it clear what they do, and they respect conventions like robots.txt.
These residential proxies have been abused by scrapers that use random legit-looking user agents and absolutely hammer websites. What is it with these scrapers just not understanding consent? It's gross.
Re: Disrupting the largest residential proxy network
#226Earlier quoted context omitted.
No, no they really aren't, but I was thinking the "scraping industry" in the sense that that's a thing. Getting hosting in smaller datacenters is simple enough, but you may need to manage your own hardware, or VMs. Many will help you get your own IP ranges and ASN, that's going to go a long way, if you don't want to get bundled in with the bad bots. This differs obviously, but having an ASN in our case means that we…
Scraping isn’t an industry. There are legitimate and illegitimate scraping pursuits. There are lots of healthy / productive businesses in the cloud and lots of scumbags, just like any enterprise. I still have no idea about your point, by the way.
Hiding behind a residential proxy and using random user agents? Gross. Learn what consent is.
Re: Disrupting the largest residential proxy network
#227Earlier quoted context omitted.
Does malicious mean interfering with Google's business model, or does it include intrusive advertising?
Malicious here means "most people who aren't trying to argue semantics or otherwise be smartasses about it would consider it malware". That's why the example I gave is a semi-popular software the allows watching YouTube without ads without a premium subscription, i.e. at least in the case I observed, I don't believe this was weaponized against apps that interfere with their business model. As for "intrusive advertisi…
Proxying traffic is not malware, since it doesn't affect me in any way.
Re: Disrupting the largest residential proxy network
#228Earlier quoted context omitted.
Getting rid of malware is good. A private for-profit company exercising its power over the Internet, not so much. We should have appropriate organizations for this.
The proxies is the reason why you get spam in your Google search result, spam in your Play store (by means of fake good reviews), basically spam in anything user generated. It directly affects Google and you, I don’t see why they should not do this.
Re: Disrupting the largest residential proxy network
#229Earlier quoted context omitted.
Scraping isn’t an industry. There are legitimate and illegitimate scraping pursuits. There are lots of healthy / productive businesses in the cloud and lots of scumbags, just like any enterprise. I still have no idea about your point, by the way.
If you have a "legitimate scraping pursuit", identify yourself appropriately that way. I'm happy to let most well-behaved scrapers access my content. Hiding behind a residential proxy and using random user agents? Gross. Learn what consent is.
You're thinking about the case of big AI companies crawling your blog. I'm talking about a small startup trying to do traditional indexing and needing to run from residential proxy to make it work.
Re: Disrupting the largest residential proxy network
#230Earlier quoted context omitted.
> Malware in random apps running on your device without your knowledge is bad. And ones that have all the indicators of compromise of Russia, Iran, DPRK, PRC, etc
Am I the only one cynically thinking that "Russia, Iran, DPRK, PRC, etc" is the "But think of the chiiildren!!!" excuse for doing this? And when Google say "IPIDEA’s proxy infrastructure is a little-known component of the digital ecosystem leveraged by a wide array of bad actors." What they really mean is " ... leveraged by actors indiscriminately scraping the web and ignoring copyright - that are not us." I can't he…