Live data from Hacker News

Disrupting the largest residential proxy network

cloud.google.com

141–150 of 230 posts

Re: Disrupting the largest residential proxy network

#141
post #35

Earlier quoted context omitted.

2k IPs is not enough to do most enterprise scale scraping. Starlink's entire ASN doesn't seem to have enough V4 addresses to handle it even.

The actual secret is to use IPv6 with varied source IPs in the same subnet, you get an insane number of IPs and 90% of anti-scraping software is not specialized enough to realize that any IP in a /64 is the same as a single IP in a /32 in IPv4.

> any IP in a /64 is the same as a single IP in a /32 in IPv4

This is very commonly true but sadly not 100%. I am suffering from a shared /64 on which a VPS is, and where other folks have sent out spam - so no more SMTP for me.

Re: Disrupting the largest residential proxy network

#142

Residential proxies are the only way to crawl and scrape. It's ironic for this article to come from the biggest scraping company that ever existed! If you crawl at 1Hz per crawled IP, no reasonable server would suffer from this. It's the few bad apples (impatient people who don't rate limit) who ruin the internet for both users and hosters alike. And then there's Google.

One thing about Google is that many anti-scraping services explicitly allow access to Google and maybe couple of other search engines. Everybody else gets to enjoy CloudFlare captcha, even when doing crawling at reasonable speeds. Rules For Thee but Not for Me

You say this like robots.txt doesn't exist.

Re: Disrupting the largest residential proxy network

#143

Earlier quoted context omitted.

Yeah, it serves the purpose of blocking this kind of proxy traffic that isn't in Google's personal best interests. Only Google is allowed to scrape the web.

"Only Google is allowed to scrape the web." If I'm not mistaken, the plaintiffs in the US v Google antitrust litigation in the DC Circuit tried to argue that website operators are biased toward allowing Google to crawl and against allowing other search engines to do the same The Court rejected this argument because the plaintiffs did not present any evidence to support it For someone who does not follow the web's his…

> For someone who does not follow the web's history, how would one produce direct evidence that the bias exists

Take a bunch of websites, fetch their robots.txt file and check how many allow GoogleBot but not others?

Re: Disrupting the largest residential proxy network

#144

Earlier quoted context omitted.

Google does not use residential proxies. This does nothing against your ability to scrape the web the Google way, AKA from your own assigned IP range, obeying robots.txt, and with an user agent that explicitly says what you're doing and gives website owners a way to opt out. What Google doesn't want (and I don't think that's a bad thing) is competitors scraping the web in bad faith, without disclosing what they're do…

> This does nothing against your ability to scrape the web the Google way I thought that Google has access to significant portions of the internet that non-Google bots won’t have access to?

Their crawler has known IPs that get a white-glove treatment by every site with a paywall for example

Re: Disrupting the largest residential proxy network

#145

It’s interesting that when Luminati, an Israeli company, does this, it’s fine. When the Chinese do this? Very bad.

They are both bad. You are showing your own bias.

No, he is referencing Google going after the Chinese company, not the Israel based one. That does not mean there is bias with the commenter at all, just that the companies operate differently and are treated differently. The country of origin is important as Israel based companies are more integrated into the western business world, and tend to at least try to show an effort in keeping spam and other things off their platforms. Now I do agree that they are both bad companies that should not be allowed to operate the way they do. I would say the same thing about the other 1000 scrapers hitting websites everyday as well (including Google).

What they did not comment directly on, is how many apps / games they might have actually removed from the Playstore with the removal of the SDKs, which would be the actual interesting data.

Re: Disrupting the largest residential proxy network

#146
post #52

Earlier quoted context omitted.

They provide an SDK for mobile developers. Here is a video of how it works. [0] They don't even hide it. [0] https://www.youtube.com/watch?v=1a9HLrwvUO4&t=15s

Of course they're pitching it like everything's above board, but from the article: > While many residential proxy providers state that they source their IP addresses ethically, our analysis shows these claims are often incorrect or overstated. Many of the malicious applications we analyzed in our investigation did not disclose that they enrolled devices into the IPIDEA proxy network. Researchers have previously found…

I love how its the "evil" Open Source project devices, and "other app stores" that are the problem, not the 100s of spyware ridden crap that is available for download from the Play store. Would be interesting to know how many copies of the SDK was found and removed from their own platform.

Re: Disrupting the largest residential proxy network

#147
post #122

Earlier quoted context omitted.

Users are OK with acting as proxies because they don't understand all the shady stuff their proxy is being used for. Also consumer ISPs generally ban this.

You could say the same about google’s terms of service.

A thousand times yes.

Re: Disrupting the largest residential proxy network

#148

Earlier quoted context omitted.

> Malware in random apps running on your device without your knowledge is bad. And ones that have all the indicators of compromise of Russia, Iran, DPRK, PRC, etc

Am I the only one cynically thinking that "Russia, Iran, DPRK, PRC, etc" is the "But think of the chiiildren!!!" excuse for doing this? And when Google say "IPIDEA’s proxy infrastructure is a little-known component of the digital ecosystem leveraged by a wide array of bad actors." What they really mean is " ... leveraged by actors indiscriminately scraping the web and ignoring copyright - that are not us." I can't he…

No, what they're saying is what they said, what you're implying reveals a strange bias. Web scraping through residential proxies? Please think through your thoughts more. There's much more effective and efficient ways to do so. Multiple bad actors, like ransomware affiliates, have been caught using residential proxy networks. But by all means, don't let facts and cyber threat intelligence get in the way.

Re: Disrupting the largest residential proxy network

#150

Earlier quoted context omitted.

> Malware in random apps running on your device without your knowledge is bad. And ones that have all the indicators of compromise of Russia, Iran, DPRK, PRC, etc

Am I the only one cynically thinking that "Russia, Iran, DPRK, PRC, etc" is the "But think of the chiiildren!!!" excuse for doing this? And when Google say "IPIDEA’s proxy infrastructure is a little-known component of the digital ecosystem leveraged by a wide array of bad actors." What they really mean is " ... leveraged by actors indiscriminately scraping the web and ignoring copyright - that are not us." I can't he…

>I can't help but feel this is just Google trying to pull the ladder up behind then and make it more difficult for other companies to collect training data.

I can very easily see this as being Google's reasoning for these actions, but let's not pretend that clandestine residential proxies aren't used for nefarious things. The vast majority of social media networks will ban - or more generally and insiously - shadow ban accounts/IPs that use known proxy IPs. This means that they are gating access to their platforms behind residential IPs (on top of their other various blackboxes and heuristics like fingerprinting). Operators of bot networks thus rely on residential proxy services to engage in their work, which ranges from mundane things like engagement farming to outright dangerous things like political astroturfing, sentiment manipulation, and propaganda dissemination.

LLMs and generative image and video models have made the creation of biased and convincing content trivial and cheap, if not free. The days of "troll farms" is over, and now the greatest expense for a bad actor wishing to influence the world with fake engagement and biased opinions is their access to platforms, which means accounts and internet connections that aren't blacklisted or shadow banned. Account maturity and reputation farming is also feeling a massive boon due to these tools, but as an independent market it also similarly requires internet connections that aren't blacklisted or shadow banned. Residential proxies are the bottleneck for the vast majority of bad actors.

Post reply on HN