Live data from Hacker News

An update on residential proxies and the scraper situation

lwn.net

311–320 of 422 posts

Re: An update on residential proxies and the scraper situation

#311

The article at the end talks about how is very easy for arbitrary apps from app stores can install a residential proxy on your phone. 10 years ago, apps had to explicitly state if they needed network access. And then the powers that be decided that really all apps need network access no matter what. And both ios and android make it hard to deny apps network access. But really, this finally explains the hordes of real…

The Bright Data “free” VPN they’re talking about requires the user to go through steps to enable it.

These aren’t as simple as downloading a free game and then the phone is compromised as long as it’s installed.

The users who install these things don’t care about permissions prompts. They’ll follow instructions to tap any prompt the instructions ask. They want the free thing and don’t care what they have to do to get it.

Re: An update on residential proxies and the scraper situation

#312

The article at the end talks about how is very easy for arbitrary apps from app stores can install a residential proxy on your phone. 10 years ago, apps had to explicitly state if they needed network access. And then the powers that be decided that really all apps need network access no matter what. And both ios and android make it hard to deny apps network access. But really, this finally explains the hordes of real…

The Bright Data “free” VPN they’re talking about requires the user to go through steps to enable it. These aren’t as simple as downloading a free game and then the phone is compromised as long as it’s installed. The users who install these things don’t care about permissions prompts. They’ll follow instructions to tap any prompt the instructions ask. They want the free thing and don’t care what they have to do to get…

The article also talks about NetNut, which was embed in many apps released into the app stores.

Re: An update on residential proxies and the scraper situation

#313
post #261
post #45

Earlier quoted context omitted.

What really confuses me is ... people always say, it's because companies are gathering data for AI training. Then why would they need to scrape the same page thousands of times per day? Edit: the article says millions of times per hour? (!?) The article is also astonished by this, and speculates it might be some kind of underground AI labs but... millions of them? Or does it only take one with too much money and a ba…

I had meta's crawler hitting the pi searcher at something like 5qps for days on end, just ... querying for substrings of pi, ignoring robots.txt, etc. it wasn't enough to break anything but it triggered a lot of alerts. I can imagine that sites with dynamic content and potentially unbounded query types or pathnames are in danger from particularly stupid crawlers.

Meta has hit me with over 1,000 requests per second. Luckily it tripped my rate limiting but geez.

Re: An update on residential proxies and the scraper situation

#314

The article at the end talks about how is very easy for arbitrary apps from app stores can install a residential proxy on your phone. 10 years ago, apps had to explicitly state if they needed network access. And then the powers that be decided that really all apps need network access no matter what. And both ios and android make it hard to deny apps network access. But really, this finally explains the hordes of real…

GrapheneOS allows you to deny network access per app pretty trivially. Google Play services make it a bit more difficult because the app might marshall the network request through that; I'm not sure how to verify that behavior when it happens.

Thanks for mentioning GrapheneOS. I'll just explain to others what "pretty trivially" means here. GrapheneOS adds a huge checkmark "Network access?" when you install an app. It's impossible to miss.

Re: An update on residential proxies and the scraper situation

#315

In theory, proof of work that is used to mine a cryptocurrency could be a solution. Bitcoin and others are already secured via massive pow computations. If we could shift that into browsers, no additional energy would be used and we could solve an issue that has been unsolved for too long: How to pay websites that provide useful information other than with ads. The question is which resources typical consumer hardwar…

Better skip the PoW part as it'll be wildly inefficient for most of the work done.

Instead, exchange web traffic for actual $. Say, some kind of tokens that are easily turned back into hard cash through a 3rd party.

Requesting a 100KB file? Okay, that'll be a $0.00002 token, please! (visitor's user agent provides it in a manner transparent to regular web users). Requesting a 3MB image? Okay, that'll be a $0.0005 token, please!

Result: niche websites earn hard cash. It doesn't matter much if you're hammered as long as the hammering comes with a corresponding flow of tokens (read: $). No token(s)? No service.

Regular web users would pay for those tokens through their normal internet service fees, and otherwise not be bothered. Massive scrapers would have to pay somehow for the tokens to be served web data at all.

In effect: put the bulk of public web sites behind a paywall. But with the bar low enough & in a manner that it's transparent for regular web users. Clicked "reload" by accident? Oops, internet service bill got upped by 2 micro-$.

Re: An update on residential proxies and the scraper situation

#316

Earlier quoted context omitted.

If it weren't a real problem, these types of articles and services wouldn't exist.

Plenty of complaints exist about things that are not real problems.

Well that's not what's happening here, lol.

Re: An update on residential proxies and the scraper situation

#317
post #77

>We have not gone with tools like Anubis, partly because it causes annoying delays for those trying to get to the site, but also partly because it seems inevitable that the scrapers will eventually find their way around it. Indeed, there are some indications that is already happening. A proof-of-work requirement is not a huge obstacle when you have millions of other people's machines to do the work on. The first argu…

Residential proxy users don't have the ability to run compute on their proxies.

https://salad.com/salad-gateway-service

https://salad.com/earn

Re: An update on residential proxies and the scraper situation

#318

Earlier quoted context omitted.

No. When someone says you’re not welcome, you’re not welcome. Regardless of the reason.

If I get kicked out of Epstein island because I refuse to **** a child, that doesn't make me a bad actor.

And then going there 1 million times under fake identities? Yeah, I'm sure that's not a bad actor.

Re: An update on residential proxies and the scraper situation

#319
post #299
post #241

Earlier quoted context omitted.

LI've emailed there in May and June about different topics and gotten no reply to a request to confirm receipt. Is there a backup method for when your algorithm throws people's email away without even informing the sender with a delivery failure notification?

We get thousands of emails - far too many to even read them all, let alone respond the way we would like to. There are other reasons why we might not respond, but overload is the main one. Edit: I only see one email in the archive related to your account. It was from Sept 2024 and we responded to it. Are you talking about a different account?

Oh, okay I didn't know you weren't getting around to all incoming emails, that makes sense then. Previously the responses were fast and reliable and I thought I was being dropped by a silent bot/spam detection algorithm (as is becoming frustratingly common).

It's indeed from two different email addresses from the one on this account. They're not that important so never mind, thanks for checking and for the reply!

Re: An update on residential proxies and the scraper situation

#320

Earlier quoted context omitted.

GrapheneOS allows you to deny network access per app pretty trivially. Google Play services make it a bit more difficult because the app might marshall the network request through that; I'm not sure how to verify that behavior when it happens.

Thanks for mentioning GrapheneOS. I'll just explain to others what "pretty trivially" means here. GrapheneOS adds a huge checkmark "Network access?" when you install an app. It's impossible to miss.

The problem is the yes/no options - I want to give access to specific endpoints - but now that’s “all of Cloudflare for everything” which means the web entire.

It used to be you could let Onkyo App™ access the five IPs for onkyo.com and be done with it.

Post reply on HN