An update on residential proxies and the scraper situation
1–10 of 422 posts
Re: An update on residential proxies and the scraper situation
#2The question is more about why the US and others can't properly enforce the bullshit all this amounts to.
Re: An update on residential proxies and the scraper situation
#3Re: An update on residential proxies and the scraper situation
#4The poison gets better every day, and the community is continuously growing. Poison Fountain, alone, transmits hundreds of gigabytes of poison per day, which goes into scrapers, git repositories on every hosting platform, social media, etc.
Part of the poisoning community on Reddit, for example: https://www.reddit.com/r/PoisonFountain/comments/1uocaii/a_n...
Re: An update on residential proxies and the scraper situation
#5I worry a lot of the anti scraping rhetoric will just injure the open web and put somebody like cloudflare in charge.
Re: An update on residential proxies and the scraper situation
#6Re: An update on residential proxies and the scraper situation
#7There is a large community of people that poison scrapers. The poison gets better every day, and the community is continuously growing. Poison Fountain, alone, transmits hundreds of gigabytes of poison per day, which goes into scrapers, git repositories on every hosting platform, social media, etc. Part of the poisoning community on Reddit, for example: https://www.reddit.com/r/PoisonFountain/comments/1uocaii/a_n...
Re: An update on residential proxies and the scraper situation
#8I think that nobody would care if I use wget or curl for few pages, e.g. because I would like to read a site as offline or archive it.
Btw average age of any page is 10 years. Deletion or structural change after acquisition is common, Signal vs Noise site recent wipe out could serve as an example why we need to archive sites.
Re: An update on residential proxies and the scraper situation
#9I wonder how much of this is traffic caused by peoples agents using web tools causing searches and fetches rather than general trawls of the internet.