Cloudflare crawl endpoint
11–20 of 194 posts
Re: Cloudflare crawl endpoint
#12Is cloudflare becoming a mob outfit? Because they are selling scraping countermeasures but are now selling scraping too. And they can pull it off because of their reach over the internet with the free DNS.
Re: Cloudflare crawl endpoint
#13Is cloudflare becoming a mob outfit? Because they are selling scraping countermeasures but are now selling scraping too. And they can pull it off because of their reach over the internet with the free DNS.
Re: Cloudflare crawl endpoint
#14Does this bypass their own anti-AI crawl measures? I'll need to test it out, especially with the labyrinth.
We're creating an internet that is becoming self-reinforcing for those who already have power and harder for anyone else. As crawling becomes difficult and expensive, only those with previously collected datasets get to play. I certainly understand individual sites wanting to limit access, but it seems unlikely that they're limiting access to the big players - and maybe even helping them since others won't be able to compete as well.
Re: Cloudflare crawl endpoint
#15I'm surprised that Cloudflare hasn't started hosting a pre-scraped version of websites that use Cloudflare's proxy - something like https://www.example.com/cdn-cgi/cached-contents.json They already have the website content in their cache, so why not just cut out the middle man of scraping services and API's like this and publish it? Obviously there's good reasons NOT to, but I am surprised they haven't started offeri…
It's entirely possible that they're doing this under the hood for cases where they can clearly identify the content they have cached is public.
Re: Cloudflare crawl endpoint
#16Does this bypass their own anti-AI crawl measures? I'll need to test it out, especially with the labyrinth.
Further down they also mention that the requests come from CFs ASN and are branded with identifying headers, so third party filters could easily block them too if they're so inclined. Seems reasonable enough.
Re: Cloudflare crawl endpoint
#17Is cloudflare becoming a mob outfit? Because they are selling scraping countermeasures but are now selling scraping too. And they can pull it off because of their reach over the internet with the free DNS.
> The /crawl endpoint respects the directives of robots.txt files, including crawl-delay. All URLs that /crawl is directed not to crawl are listed in the response with "status": "disallowed".
You don't need any scraping countermeasures for crawlers like those.
Re: Cloudflare crawl endpoint
#18Is cloudflare becoming a mob outfit? Because they are selling scraping countermeasures but are now selling scraping too. And they can pull it off because of their reach over the internet with the free DNS.
The fact that 30%+ of the web relies on their caching services, routablility services and DDoS protection services is the main pull.
Their DNS is only really for data collection and to front as "good will"
Re: Cloudflare crawl endpoint
#19Re: Cloudflare crawl endpoint
#20This might be really great! I had the idea after buying https://mirror.forum recently (which I talked in discord and archiveteam irc servers) that I wanted to preserve/mirror forums (especially tech) related [Think TinyCoreLinux] since Archive.org is really really great but I would prefer some other efforts as well within this space. I didn't want to scrape/crawl it myself because I felt like it would feel like yet a…
You feel better paying someone to do the same thimg?