Facebook's Fascination with My Robots.txt
blog.nytsoi.net
Facebook's Fascination with My Robots.txt
1–10 of 55 posts
Re: Facebook's Fascination with My Robots.txt
#2Re: Facebook's Fascination with My Robots.txt
#3Re: Facebook's Fascination with My Robots.txt
#4Re: Facebook's Fascination with My Robots.txt
#5If you've been in any big company you'll know things perpetually run in a degraded, somewhat broken mode. They've even made up the term "error budget" because they can't be bothered to fix the broken shit so now there's an acceptable level of brokenness.
Re: Facebook's Fascination with My Robots.txt
#6Did you try adding a Cache-Control response header?
Is this where all that hardware for AI projects is going? To data centers that just uncritically hits the same URL over and over without checking if the content of a site or page has chanced since the last visit then and calculate a proper retry interval. Search engine crawlers 25 - 30 years ago could do this.
Hit the URL once per day, if it chances daily, try twice a day. If it hasn't chanced in a week, maybe only retry twice per week.
Re: Facebook's Fascination with My Robots.txt
#7Re: Facebook's Fascination with My Robots.txt
#8Did you try adding a Cache-Control response header?
Even if they haven't added any cache control headers, what kind a of lazy Meta engineer designed their crawler with to just pull the same URL multiple times a second? Is this where all that hardware for AI projects is going? To data centers that just uncritically hits the same URL over and over without checking if the content of a site or page has chanced since the last visit then and calculate a proper retry interva…
Re: Facebook's Fascination with My Robots.txt
#9Re: Facebook's Fascination with My Robots.txt
#10Have you considered serving a zip bomb to this user agent?