If I use Cloudflare, what can I do — if anything — to avoid disruption when they go down?
Cloudflare outage on June 21, 2022
71–80 of 234 posts
Re: Cloudflare outage on June 21, 2022
#72Re: Cloudflare outage on June 21, 2022
#73Yet another BGP caused outage. At some point we should collect all of them: - Cloudflare 2022 (this one) - Facebook 2021: https://news.ycombinator.com/item?id=28752131 - this one probably had the single biggest impact, since engineers got locked out of their systems, which made the fixing part look like a sci-fi movie - (Indirectly caused by BGP: Cloudflare 2020: https://blog.cloudflare.com/cloudflare-outage-on-july-…
Re: Cloudflare outage on June 21, 2022
#74Yet another BGP caused outage. At some point we should collect all of them: - Cloudflare 2022 (this one) - Facebook 2021: https://news.ycombinator.com/item?id=28752131 - this one probably had the single biggest impact, since engineers got locked out of their systems, which made the fixing part look like a sci-fi movie - (Indirectly caused by BGP: Cloudflare 2020: https://blog.cloudflare.com/cloudflare-outage-on-july-…
The internet runs on BGP, I would think that most internet issues would be a result of BGP then.
It just seems that most of these are local enough and the Internet resilient enough that they don't cause global issues. Maybe the exception would be AWS us-east-1 outages :-)
Re: Cloudflare outage on June 21, 2022
#75Looking at the hourly chart in Google Analytics (compared to the previous day) there isn't even a blip during this outage.
So for all the advantages we get from Cloudflare (caching, WAF, security [our WP admin is secured with Cloudflare Teams], redirects, page rules, etc) I'll take these minor outages that make HN go apeshit.
Of course it helped that most our traffic is from the US and this happened when it did but in the past week alone we served over 180 countries which Cloudflare helps make sure is nice and fast :D
Re: Cloudflare outage on June 21, 2022
#76TODO: use commit-confirm for automated rollbacks Sounds like a good idea!
Re: Cloudflare outage on June 21, 2022
#77In a world where it can take weeks for other companies to publish a postmortem after an outage (if they ever do), I never ceases to amaze me how quickly CF manage to get something like this out. I think it's a testament to their Ops/Incident response teams and internal processes, it builds confidence in their ability to respond quickly when something does go wrong. Incredible work!
I agree, I think the transparency builds trust and I encourage it where I can. The counter thought I had when reading this case though, is it almost feels too fast. What I mean by that is I hope there isn't an incentive to wrap up the internal investigation quickly and write the blog and send it, and go we're done. Doing incident response (both outage and security), the tactical fixes for a specific problem are usual…
> This was delayed as network engineers walked over each other's changes, reverting the previous reverts, causing the problem to re-appear sporadically.
They are running as fast as they can and this extended the incident. There is a “slow is smooth, smooth is fast” lesson in here. I’d rather have a team that takes a day to put up the blog post, but doesn’t unnecessarily extend downtime because they are sprinting.
Re: Cloudflare outage on June 21, 2022
#78Re: Cloudflare outage on June 21, 2022
#79Earlier quoted context omitted.
I feel like others lose opportunities by not doing the same. By publishing early and publishing the details they: keep the company in the news with positive stuff (free ad), get an internal documentation of the incident (ignoring the customer oriented "we're sorry" part), effectively get a free recruitment post (you're reading this because you're in tech and we do cool stuff, wink), release some internal architecture…
> I wonder how much those posts are calculated and how much organic/culture related. Don't companies have a fiduciary duty to calculate things; the reason for doing something actually cannot just be that it's a nice thing to do? Not down to the word, but at least the general decision to be this way?
Re: Cloudflare outage on June 21, 2022
#80Earlier quoted context omitted.
To further add to your point, the CTO is the one who shared it here & the CEO is incredibly active on forums & social media everywhere with customers. Communication has always been one of their strengths.
I do wonder what would happen should happen if either of them left the company, I feel like there's a lot of trust on HN (and other places) that's heavily attached to them as individuals and their track record of good communication.
But once it happens things will change and, to be honest, likely for the worse.
edit> fix typo