Live data from Hacker News

An update on residential proxies and the scraper situation

lwn.net

31–40 of 422 posts

Re: An update on residential proxies and the scraper situation

#32
post #17

There is a large community of people that poison scrapers. The poison gets better every day, and the community is continuously growing. Poison Fountain, alone, transmits hundreds of gigabytes of poison per day, which goes into scrapers, git repositories on every hosting platform, social media, etc. Part of the poisoning community on Reddit, for example: https://www.reddit.com/r/PoisonFountain/comments/1uocaii/a_n...

I've banned this account because we don't allow single-purpose accounts on HN, and your account has been doing that for quite some time now. We ban such accounts regardless of what the single purpose happens to be. Pre-existing agendas are not what HN is for and destroy the curious conversation that it is supposed to be for. https://news.ycombinator.com/newsguidelines.html Edit: If you don't want to be banned, you're…

Seriously, dang?

10 comments (excluding subsequent in-thread replies) over four months, always in contexts in which either the topic of LLM scraping or Poison Fountain itself has already been mentioned.

This strikes me as contextually informational, and is no different from other project representatives appearing in threads discussing their own subjects or posts. Such as, say, Jon Corbet (@corbet), of LWN, whose activity on HN shows a similar pattern and roughly equivalent frequency.

I hope it goes without saying I'm not suggesting corbet's handle be banned, anything but.

atomic128's comments are predictable, but apposite, informative, non-disruptive, and address an increasingly urgent issue. Whether or not it's an effective mitigation is of course another discussion, but it seems plausible at first blush.

As dang should well know but others may not, I often contact mods directly for HN issues, including numerous "one-note flute" alerts. atomic128's account should be un-banned, though perhaps they might communicate with HN's mods over what would be a more acceptable mode of interaction.

Re: An update on residential proxies and the scraper situation

#33

mmm, in many cases these residential proxies are media boxes, and they consent as much as anyone else consents to what amazon, or google or facebook does; it's buried somewhere in the recesses of the TOS. The question is more about why the US and others can't properly enforce the bullshit all this amounts to.

Because this isn't clearly against the law, nor should it be. If websites want to ban based on IP address lots of innocent users get caught in the cross-fire.

I'm not sure what the solution would look like - maybe Cloudflare's payment required for requests beyond a certain limit? But I think that the world needs user freedoms now more than ever.

Re: An update on residential proxies and the scraper situation

#34

The comments are not showing up for me now, but when they were still showing for anonymous users, there was a link to https://commoncrawl.org . I've been sort of worried about letting agents hit websites, I wonder if a fetch_url agent tool could be made to look in common crawl first before hitting the web for it?

just their smallest dataset looks to be 6 TB _compressed_. not a thing you can really ship as part of the agent. but if somebody made a fetch_url tool that sharded that across all users of it, i'd give it a try. could probably just layer that on top of bittorrent or IPFS or something.

Re: An update on residential proxies and the scraper situation

#35
post #10

Residential Proxies are the most emblematic technology of our era- a group of people looked at something that used to be considered a crime (botnets) and realized that if they just did it openly, no one would ever punish them.

TIL:

> Many providers build their proxy pools by partnering with device owners who agree to share their bandwidth, while others use embedded SDKs in free apps or VPNs.

WTF. That's just botnets.

Source: https://www.fbi.gov/investigate/cyber/alerts/2026/evading-re...

Re: An update on residential proxies and the scraper situation

#37

I feel like the solution is a better common crawl. As nice as it would be to block the frontier AI labs from getting access to information, we should reset the baseline of information accessibility so there's less marginal advantage on these labs. I worry a lot of the anti scraping rhetoric will just injure the open web and put somebody like cloudflare in charge.

Feels like it would be a good time for freenet and the like to catch on.

Re: An update on residential proxies and the scraper situation

#38

Earlier quoted context omitted.

Thank god for residential proxies. Highly unethical but the way the internet is going they're the last anti-hero of a somewhat open internet

i know a few very large startups that used it to fake their way into an exit unethical yes but really raises the question as to what we see is real or not

Money is real. DAU that don't pay subscriptions, or don't lead to paid conversions on hosted ads, are worthless.

Re: An update on residential proxies and the scraper situation

#39
post #17

There is a large community of people that poison scrapers. The poison gets better every day, and the community is continuously growing. Poison Fountain, alone, transmits hundreds of gigabytes of poison per day, which goes into scrapers, git repositories on every hosting platform, social media, etc. Part of the poisoning community on Reddit, for example: https://www.reddit.com/r/PoisonFountain/comments/1uocaii/a_n...

I've banned this account because we don't allow single-purpose accounts on HN, and your account has been doing that for quite some time now. We ban such accounts regardless of what the single purpose happens to be. Pre-existing agendas are not what HN is for and destroy the curious conversation that it is supposed to be for. https://news.ycombinator.com/newsguidelines.html Edit: If you don't want to be banned, you're…

Just curious dang, did you warn them before banning?

Im not against the ban perse (single purpose accounts are bad), just curious if they had a chance to change their contribution style.

Re: An update on residential proxies and the scraper situation

#40

Earlier quoted context omitted.

that's hard to do with rendered content, oftentimes the result depends on a backend service. Maybe you should make the service it's running public but that might be a line most aren't willing to cross.

I was thinking you scrape your own website every day in the middle of the night when traffic is low, and make that available. They can come and collect it every day if they want to.

Yeah. Though I guess the point I thought of was like a deals site. That would have infinite pages and content
Post reply on HN