Turn entire websites into LLM-ready data
firecrawl.dev
Turn entire websites into LLM-ready data
1–10 of 16 posts
Re: Turn entire websites into LLM-ready data
#2This is cool, seems like a nice way to easily add context for ChatGPT at the least
Re: Turn entire websites into LLM-ready data
#3What user agent does it use and do you obey robots.txt?
* apparently not as I tested it with LinkedIn.com which blocks most non-google/bing bots and it crawls it OK
Re: Turn entire websites into LLM-ready data
#4[dead]
Re: Turn entire websites into LLM-ready data
#5[deleted]
Re: Turn entire websites into LLM-ready data
#6What user agent does it use and do you obey robots.txt? * apparently not as I tested it with LinkedIn.com which blocks most non-google/bing bots and it crawls it OK
> FireCrawl is built to navigate common web scraping challenges, including reverse proxies, rate limits, and caching
They probably ignore robots.txt
Re: Turn entire websites into LLM-ready data
#7This is cool, seems like a nice way to easily add context for ChatGPT at the least
* Creator here - Thats the goal!
Re: Turn entire websites into LLM-ready data
#8[deleted]
Re: Turn entire websites into LLM-ready data
#9Does this respect any anti AI related scraping rules set forth by website owners?