Live data from Hacker News

Disrupting the largest residential proxy network

cloud.google.com

171–180 of 230 posts

Re: Disrupting the largest residential proxy network

#171

Earlier quoted context omitted.

>I can't help but feel this is just Google trying to pull the ladder up behind then and make it more difficult for other companies to collect training data. I can very easily see this as being Google's reasoning for these actions, but let's not pretend that clandestine residential proxies aren't used for nefarious things. The vast majority of social media networks will ban - or more generally and insiously - shadow b…

> The vast majority of social media networks will ban - or more generally and insiously - shadow ban accounts/IPs that use known proxy IPs. This means that they are gating access to their platforms behind residential IPs (on top of their other various blackboxes and heuristics like fingerprinting) Social media will ban proxy IPs, yet gleefully force you to provide your ID if you happen to connect from the wrong patch…

> The only way somebody living in the UK can access Imgur is through a residential proxy.

And very little of value was lost.

> This really just sounds like a rehash of the argument against encryption. "Bad people use it, so it should go away" - never mind that there are completely legitimate uses for it.

Except that almost everything that uses encryption has some legitimate use. There are pretty much no legitimate uses for residential proxies, and their use in flooding the Internet with crap greatly outweighs that.

If I plumbed a 30cm sewage line straight into your living room would you be happy with it? Okay, well, tell you what, let's make it totally legit - I'll drop a tasty ripe strawberry into the stream of effluent every so often, how about that?

Re: Disrupting the largest residential proxy network

#172

Residential proxies are the only way to crawl and scrape. It's ironic for this article to come from the biggest scraping company that ever existed! If you crawl at 1Hz per crawled IP, no reasonable server would suffer from this. It's the few bad apples (impatient people who don't rate limit) who ruin the internet for both users and hosters alike. And then there's Google.

One thing about Google is that many anti-scraping services explicitly allow access to Google and maybe couple of other search engines. Everybody else gets to enjoy CloudFlare captcha, even when doing crawling at reasonable speeds. Rules For Thee but Not for Me

Why are you scraping sites in the first place? What legitimate reason is there for you doing that?

Re: Disrupting the largest residential proxy network

#173
post #59

Earlier quoted context omitted.

I've tried it, and my account was shadowbanned a few hours after I created it. It's very obnoxious.

Reddit bots shadowban almost everyone who post before they have enough comment karma. Nothing to do with Tor or VPN.

I didn't try posting, I tried commenting.

Re: Disrupting the largest residential proxy network

#174
post #41

We need more residential proxies, not less. I've had enough of companies saying "you're connecting from an AWS IP address, therefore you aren't allowed in, or must buy enterprise licensing". Reddit is an example which totally blocks all data to non-residential IP's. I want exactly the same content visible no matter who you are or where you are connecting from, and a robust network of residential proxies is a stepping…

I live in the UK and can't view a large portion of the internet without having to submit my ID to _every_ site serving anything deemed "not safe the for the children". I had a question about a new piercing and couldn't get info on it from Reddit because of that. I try using a VPN and they're blocked too. Luckily, I work at a copmany selling proxies so I've got free proxies whenever I want, but I shouldn't _need_ to u…

> I live in the UK and can't view a large portion of the internet without having to submit my ID to _every_ site serving anything deemed "not safe the for the children".

Really? Because I live in the UK and I've never been asked for my ID for anything.

Re: Disrupting the largest residential proxy network

#176

Earlier quoted context omitted.

> Malware in random apps running on your device without your knowledge is bad. And ones that have all the indicators of compromise of Russia, Iran, DPRK, PRC, etc

Am I the only one cynically thinking that "Russia, Iran, DPRK, PRC, etc" is the "But think of the chiiildren!!!" excuse for doing this? And when Google say "IPIDEA’s proxy infrastructure is a little-known component of the digital ecosystem leveraged by a wide array of bad actors." What they really mean is " ... leveraged by actors indiscriminately scraping the web and ignoring copyright - that are not us." I can't he…

> Am I the only one cynically thinking that "Russia, Iran, DPRK, PRC, etc" is the "But think of the chiiildren!!!" excuse for doing this?

Maybe. But until I dropped all traffic from pretty much every mobile network provider in Russia and Israel, I'd get up every morning to a couple of thousand new users of whom a couple of hundred had consistently within a few hundred milliseconds created an account, clicked on the activation link, and then posted a bunch of messages in every forum category spreading hate speech.

Re: Disrupting the largest residential proxy network

#177

We need more residential proxies, not less. I've had enough of companies saying "you're connecting from an AWS IP address, therefore you aren't allowed in, or must buy enterprise licensing". Reddit is an example which totally blocks all data to non-residential IP's. I want exactly the same content visible no matter who you are or where you are connecting from, and a robust network of residential proxies is a stepping…

> I've had enough of companies saying "you're connecting from an AWS IP address I run a honeypot and the amount of bot traffic coming from AWS is insane. It's like 80% before filtering, and it's 100% illegitimate.

> it's 100% illegitimate.

Based on what?

I think perhaps you merely meant to say that more than 99% of it is illegitimate?

Re: Disrupting the largest residential proxy network

#178

Earlier quoted context omitted.

Yeah, it serves the purpose of blocking this kind of proxy traffic that isn't in Google's personal best interests. Only Google is allowed to scrape the web.

"Only Google is allowed to scrape the web." If I'm not mistaken, the plaintiffs in the US v Google antitrust litigation in the DC Circuit tried to argue that website operators are biased toward allowing Google to crawl and against allowing other search engines to do the same The Court rejected this argument because the plaintiffs did not present any evidence to support it For someone who does not follow the web's his…

Common Crawl provides gzipped robots.txt collections

Re: Disrupting the largest residential proxy network

#179

Residential proxies are the only way to crawl and scrape. It's ironic for this article to come from the biggest scraping company that ever existed! If you crawl at 1Hz per crawled IP, no reasonable server would suffer from this. It's the few bad apples (impatient people who don't rate limit) who ruin the internet for both users and hosters alike. And then there's Google.

do we think a scraper should be allowed to take whatever means necessary to scrape a site if that site explicitly denies that scraper access?

if someone is abusing my site, and i block them in an attempt to stop that abuse, do we think that they are correct to tell me it doesn’t matter what i think and to use any methods they want to keep abusing it?

that seems wrong to me.

Re: Disrupting the largest residential proxy network

#180

Earlier quoted context omitted.

One thing about Google is that many anti-scraping services explicitly allow access to Google and maybe couple of other search engines. Everybody else gets to enjoy CloudFlare captcha, even when doing crawling at reasonable speeds. Rules For Thee but Not for Me

You say this like robots.txt doesn't exist.

it almost sounds like they’re saying the contents of robots.txt shouldn’t matter… because google exists? or something?

implying “robots.txt explicitly says i can’t scrape their site, well i want that data, so im directing my bot to take it anyway.”

Post reply on HN