Live data from Hacker News

An update on residential proxies and the scraper situation

lwn.net

51–60 of 422 posts

Re: An update on residential proxies and the scraper situation

#51

mmm, in many cases these residential proxies are media boxes, and they consent as much as anyone else consents to what amazon, or google or facebook does; it's buried somewhere in the recesses of the TOS. The question is more about why the US and others can't properly enforce the bullshit all this amounts to.

What exactly should be illegal here? Scraping websites? AI agents? Not following robots.txt?

Re: An update on residential proxies and the scraper situation

#53
post #31

One article mentioned in the OP was discussed here: Disrupting the largest residential proxy network - https://news.ycombinator.com/item?id=46802748 - Jan 2026 (221 comments)

How does HN fare with scraper load? Is it just CDN and pay the extra bandwidth bill for anon hit requests?

For one datapoint ...

I have a custom HN CSS which includes some formatting of different sets of user accounts. Admins, for example, get orange highlighting and a dragon emoji (for one does not meddle in the affairs of ...).

Also included are leaders, which is the one part of my CSS build script which is, or at least was until a few minutes ago, dynamic. Presently HN is returning "sorry" to my curl request. Given that I run that build manually a few times a month, it's not a matter of hitting HN with frequent scrapes. But HN has become increasingly scrape-hostile over time.

Back in 2023 I did a crawl of all of HN's front-page daily history (365.25 days/year * 17 years, so about 6,200 requests), to answer a question which had come up about what was/wasn't mentioned in submission titles. That scrape included a delay (probably either 1 or 10 seconds, possibly more, I don't recall which and may have run the fetch directly from the command line), and ran (initially) without issues. I don't think it would fly today.

I reported on findings at the time and several times since:

https://news.ycombinator.com/item?id=36078578>

https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...>

Re: An update on residential proxies and the scraper situation

#54
post #31

One article mentioned in the OP was discussed here: Disrupting the largest residential proxy network - https://news.ycombinator.com/item?id=46802748 - Jan 2026 (221 comments)

How does HN fare with scraper load? Is it just CDN and pay the extra bandwidth bill for anon hit requests?

Not dang and not the person you are asking but there is no CDN. HN is just two servers running BSD one active and one standby. HN is all text so there is not much bandwidth usage.

I did an experiment and linked from HN to my lame blog site and disabled all my anti-scraping measures. Even with all the bots I did not see that much traffic. I suspect some people are specifically being targeted by very poorly configured or very poorly written archiving scripts. Just one example thread discussing this with someone on HN [1]. Each case of being targeted will require looking at generalized characteristics but most are easy to stop in my opinion and experience.

[1] - https://news.ycombinator.com/item?id=48416693

Re: An update on residential proxies and the scraper situation

#55

Earlier quoted context omitted.

Thank god for residential proxies. Highly unethical but the way the internet is going they're the last anti-hero of a somewhat open internet

i know a few very large startups that used it to fake their way into an exit unethical yes but really raises the question as to what we see is real or not

"Raises the question of what we see is real"

No they really don't, dishonest founders do that.

You're one with the lower case shibboleth so I have no doubt you surround yourself with dishonest founders, but faking users is pretty damn low on the usecases for residential proxies.

I said they're unethical because they tend to be hidden in innocuous seeming apps or sprung on unwitting individuals via clickwraps on their smart devices.

Re: An update on residential proxies and the scraper situation

#56

Earlier quoted context omitted.

Seriously, dang? 10 comments (excluding subsequent in-thread replies) over four months, always in contexts in which either the topic of LLM scraping or Poison Fountain itself has already been mentioned. This strikes me as contextually informational, and is no different from other project representatives appearing in threads discussing their own subjects or posts. Such as, say, Jon Corbet (@corbet), of LWN, whose acti…

I think the reasoning is about having alt accounts for different purposes. He intention is to map one human to one account and have all of their thoughts from that one account, instead of one human having one account to discuss scraping on, and a different account to discuss crypto on.

I'm pretty confident it's not that.

HN's prime directive is "anything that gratifies one's intellectual curiosity": https://news.ycombinator.com/newsguidelines.html> and many, many, many dang comments.

I'm pretty sure that the specific gripe is posting excessively (not even necessarily exclusively) on a single topic or theme. See https://news.ycombinator.com/item?id=19392902> for a more detailed comment from dang.

Occasional alts are explicitly permitted, though not to engage in abuse (e.g., mutual admiration societies, sock-puppetry). See: https://news.ycombinator.com/item?id=9963551> https://news.ycombinator.com/item?id=9823379> (both against sock puppetry) and https://news.ycombinator.com/item?id=9122086> and https://news.ycombinator.com/item?id=7504621> (on where throwaways are/aren't permitted).

Where HN does favour persistent accounts the stated claim is to foster community, rather than for nefarious tracking purposes: https://news.ycombinator.com/item?id=18082346> and https://news.ycombinator.com/newsguidelines.html>. From that last:

Throwaway accounts are ok for sensitive information, but please don't create accounts routinely. HN is a community—users should have an identity that others can relate to.

Re: An update on residential proxies and the scraper situation

#57
post #10

Residential Proxies are the most emblematic technology of our era- a group of people looked at something that used to be considered a crime (botnets) and realized that if they just did it openly, no one would ever punish them.

TIL: > Many providers build their proxy pools by partnering with device owners who agree to share their bandwidth, while others use embedded SDKs in free apps or VPNs. WTF. That's just botnets. Source: https://www.fbi.gov/investigate/cyber/alerts/2026/evading-re...

Mmm... your quote (IDK where it's from) mentions them having consent from device owners, but your FBI link cautions on how to avoid getting infected by malware.

If they have consent, they're not really botnets. Botnets involve infecting devices without the owners knowing.

With consent, it wouldn't be much different from e.g. open WiFis at restaurants and hotels, companies using a single ISP and single public IPv4 address for all their employees, and most VPN services.

Re: An update on residential proxies and the scraper situation

#58
post #54

Earlier quoted context omitted.

How does HN fare with scraper load? Is it just CDN and pay the extra bandwidth bill for anon hit requests?

Not dang and not the person you are asking but there is no CDN. HN is just two servers running BSD one active and one standby . HN is all text so there is not much bandwidth usage. I did an experiment and linked from HN to my lame blog site and disabled all my anti-scraping measures. Even with all the bots I did not see that much traffic. I suspect some people are specifically being targeted by very poorly configured…

And the backup is field-tested to fall over within a few hours of the primary ;-)

https://news.ycombinator.com/item?id=32048148>

https://news.ycombinator.com/item?id=32031243>

(AFAIK that specific failure mode has in fact been addressed.)

Re: An update on residential proxies and the scraper situation

#59
post #10

Residential Proxies are the most emblematic technology of our era- a group of people looked at something that used to be considered a crime (botnets) and realized that if they just did it openly, no one would ever punish them.

Thank god for residential proxies. Highly unethical but the way the internet is going they're the last anti-hero of a somewhat open internet

By providing a way for corporate AI scrapers to operate with impunity and force the last few independently-run websites to move to the cloud?

Re: An update on residential proxies and the scraper situation

#60
post #57

Earlier quoted context omitted.

TIL: > Many providers build their proxy pools by partnering with device owners who agree to share their bandwidth, while others use embedded SDKs in free apps or VPNs. WTF. That's just botnets. Source: https://www.fbi.gov/investigate/cyber/alerts/2026/evading-re...

Mmm... your quote (IDK where it's from) mentions them having consent from device owners, but your FBI link cautions on how to avoid getting infected by malware. If they have consent, they're not really botnets. Botnets involve infecting devices without the owners knowing. With consent, it wouldn't be much different from e.g. open WiFis at restaurants and hotels, companies using a single ISP and single public IPv4 add…

by consent they mean a dialog/EULA with careful wording was display and the user clicked ok.
Post reply on HN