Show HN: robots.txt as a service, check web crawl rules through an API
1–9 of 9 posts
Re: Show HN: robots.txt as a service, check web crawl rules through an API
#2Give me some feedback!
Re: Show HN: robots.txt as a service, check web crawl rules through an API
#3It looks like a great way for you to discover URLs but like a terribly slow way for people to avoid implementing robots.txt rules.
Re: Show HN: robots.txt as a service, check web crawl rules through an API
#4I understand the unethical nature of the above method, however, I see it happening quite a lot in practice.
Re: Show HN: robots.txt as a service, check web crawl rules through an API
#5What's the plan here? Check for a sitemap.xml (which generally only contains crawlable URLs anyway) or crawl the index and look for all links and send a request to your service for every URL before crawling it?
I personally think it would be better suited as a library where you can pass it a robots.txt and it'll let you know if you can crawl a URL based on that.
Re: Show HN: robots.txt as a service, check web crawl rules through an API
#6Why a service and not a library? It looks like a great way for you to discover URLs but like a terribly slow way for people to avoid implementing robots.txt rules.
The aim of this project is only check if a given web resource can be crawled by a user-agent, but using a API
Re: Show HN: robots.txt as a service, check web crawl rules through an API
#7While this looks good, I don't think it's feasible for a web crawler in most cases. Crawlers want to crawl a ton of URLs and it would have to make a request to your service for each and every URL. What's the plan here? Check for a sitemap.xml (which generally only contains crawlable URLs anyway) or crawl the index and look for all links and send a request to your service for every URL before crawling it? I personally…
For example if you want to check the url https://example.com/test/user/1 with a user agent MyUserAgentBot, the first request can be slow (~730ms) but subsequent requests with different paths but same base url, port and protocol, will use the cached version (just ~190ms). Note that this version is in alpha and many things can be optimized. The balance between managing these files in different projects or the time between network requests must be sought.
Anyway, any person can compile the parser module and create a library to check robots.txt rules by itself ;-)
PS: thanks for the feedback
Re: Show HN: robots.txt as a service, check web crawl rules through an API
#8The service is very nice and I understand your reason for developing it. I see this service to be having more value in helping companies find all the web pages, rather than just the allowed ones. I understand the unethical nature of the above method, however, I see it happening quite a lot in practice.
Re: Show HN: robots.txt as a service, check web crawl rules through an API
#9The service is very nice and I understand your reason for developing it. I see this service to be having more value in helping companies find all the web pages, rather than just the allowed ones. I understand the unethical nature of the above method, however, I see it happening quite a lot in practice.
Yes, in the practice people sometimes don't want to be polite with webmasters, and choose not obey robots.txt rules. Thanks for the suggestion!