Earlier quoted context omitted.
What really confuses me is ... people always say, it's because companies are gathering data for AI training. Then why would they need to scrape the same page thousands of times per day? Edit: the article says millions of times per hour? (!?) The article is also astonished by this, and speculates it might be some kind of underground AI labs but... millions of them? Or does it only take one with too much money and a ba…
I always thought it's the web search tool. Grok actually shows a number of sources used for an answer. Once I asked it something simple and it apparently scanned 200 different websites. And it was just a short prompt. Now imagine millions of users asking for something multiple times a day.
An update on residential proxies and the scraper situation
331–340 of 422 posts
Re: An update on residential proxies and the scraper situation
#332> ...we have tried to minimize the impact on real readers as much as possible. We have not gone with tools like Anubis, partly because it causes annoying delays for those trying to get to the site, but also partly because it seems inevitable that the scrapers will eventually find their way around it. Indeed, there are some indications that is already happening. A proof-of-work requirement is not a huge obstacle when…
Or you can go full Reddit and just block anything that seems even remotely suspicious. Your sibling, roommate, neighbor that uses your internet, previous IP owner, posts too much? You get blocked too. Using VPN? Blocked. Your iPhone is too old, blocked. Your screen brightness too low? Believe or not, blocked.
Re: An update on residential proxies and the scraper situation
#333Earlier quoted context omitted.
Or you can go full Reddit and just block anything that seems even remotely suspicious. Your sibling, roommate, neighbor that uses your internet, previous IP owner, posts too much? You get blocked too. Using VPN? Blocked. Your iPhone is too old, blocked. Your screen brightness too low? Believe or not, blocked.
I quit reddit because of all of this nonsense. After 15 years on reddit, my life has been much better for quitting it. Reddit is a cesspool.
Re: An update on residential proxies and the scraper situation
#334> ...we have tried to minimize the impact on real readers as much as possible. We have not gone with tools like Anubis, partly because it causes annoying delays for those trying to get to the site, but also partly because it seems inevitable that the scrapers will eventually find their way around it. Indeed, there are some indications that is already happening. A proof-of-work requirement is not a huge obstacle when…
Re: An update on residential proxies and the scraper situation
#335Earlier quoted context omitted.
No one's firing up a residential proxy to read your blog, and the corporate AI scrapers have all the resources in the world even without residential proxies They're most useful for getting information from the cloud hosted sites that hoarde most of humanity's output today like Youtube and Reddit.
The Bright Data mentioned in the article, as well as other similarly malicious but even harder to identify parties, most certainly do fire up residential proxies, no matter what the site, no matter how useless or duplicate the data which they're trying to get. Not at first - they start by trying to get your content from cheap data center connections - but as soon as some kind of bot-mitigation appears, they move to r…
Re: An update on residential proxies and the scraper situation
#336Re: An update on residential proxies and the scraper situation
#337Earlier quoted context omitted.
Well this is not what is happening in practice, Wikipedia / Wikidata, OpenStreetMap, OpenFoodFacts... All provide APIs and even a full dump of their database available to download for free, but no, the stupid bots still DDoS them 24h/24.
Why don't they take legal action?
Re: An update on residential proxies and the scraper situation
#338Earlier quoted context omitted.
> There must be countless individuals all over the world who suddenly can't log into their Gmail or create any new accounts because a fraudster sent spam from their IP. Places with open WiFi like hotels and restaurants would be having the same problem. People on CGNATs would be having the same problem. An IP doesn't correspond with a single user.
Thank you. Gmail must not be like our fellow HN users we see here, quoting a couple: “I'm tiny and only run little personal stuff. I just block vast IP address blocks.” “Apologies. :( Since you say you've never visited the website before, then that means you're either in one of the countries or in one of the residential IP ranges that I've had to block.” Although Google isn’t afraid to completely block iCloud private…
Re: An update on residential proxies and the scraper situation
#339Earlier quoted context omitted.
If I get kicked out of Epstein island because I refuse to **** a child, that doesn't make me a bad actor.
And then going there 1 million times under fake identities? Yeah, I'm sure that's not a bad actor.
Either you get shown the door because you refuse to take part in unethical behavior, in which case you wouldn’t want to return anyway. Or you get shown the door because you’re simply not welcome or because YOU are the unethical one, in which case no means no.
I’m sure someone will be able to come up with some kind of edge case where this logic doesn’t work, but that doesn’t matter. This is about websites saying no to agents, bots, and scrapers. And no means no.
Re: An update on residential proxies and the scraper situation
#340Earlier quoted context omitted.
I kind of wish the recent Google monopoly court ruling had forced Google to open up their index to anyone, not just Perplexity/other big players.
That's really a huge issue right now (to some extent even before the AI hype) that almost everywhere google is effectively the only entity explicitly allowed to scrape.