What's it like to be an engineer designing and working on these systems? Must be sooo fulfiling! #Goals; Y'all are my heores!!
Cloudflare outage on June 21, 2022
51–60 of 234 posts
Re: Cloudflare outage on June 21, 2022
#52Yet another BGP caused outage. At some point we should collect all of them: - Cloudflare 2022 (this one) - Facebook 2021: https://news.ycombinator.com/item?id=28752131 - this one probably had the single biggest impact, since engineers got locked out of their systems, which made the fixing part look like a sci-fi movie - (Indirectly caused by BGP: Cloudflare 2020: https://blog.cloudflare.com/cloudflare-outage-on-july-…
Some more:
- Google 2016, configuration management bug/BGP: https://status.cloud.google.com/incident/compute/16007
- Valve 2015: https://www.thousandeyes.com/blog/steam-outage-monitor-data-...
- Cloudflare 2013: https://blog.cloudflare.com/todays-outage-post-mortem-82515/
Re: Cloudflare outage on June 21, 2022
#5307:42: The last of the reverts has been completed. This was delayed as network engineers walked over each other's changes, reverting the previous reverts, causing the problem to re-appear sporadically. Ouch
Re: Cloudflare outage on June 21, 2022
#54The problem is, we couldn't tell all our client they should change this :(
Re: Cloudflare outage on June 21, 2022
#55Really interesting that 19 cities handle 50% of the requests.
Re: Cloudflare outage on June 21, 2022
#56Re: Cloudflare outage on June 21, 2022
#57who will make the abstraction as a service we all need to protect us from config changes
Re: Cloudflare outage on June 21, 2022
#58In a world where it can take weeks for other companies to publish a postmortem after an outage (if they ever do), I never ceases to amaze me how quickly CF manage to get something like this out. I think it's a testament to their Ops/Incident response teams and internal processes, it builds confidence in their ability to respond quickly when something does go wrong. Incredible work!
I feel like others lose opportunities by not doing the same. By publishing early and publishing the details they: keep the company in the news with positive stuff (free ad), get an internal documentation of the incident (ignoring the customer oriented "we're sorry" part), effectively get a free recruitment post (you're reading this because you're in tech and we do cool stuff, wink), release some internal architecture…
Don't companies have a fiduciary duty to calculate things; the reason for doing something actually cannot just be that it's a nice thing to do? Not down to the word, but at least the general decision to be this way?
Re: Cloudflare outage on June 21, 2022
#59Re: Cloudflare outage on June 21, 2022
#60In a world where it can take weeks for other companies to publish a postmortem after an outage (if they ever do), I never ceases to amaze me how quickly CF manage to get something like this out. I think it's a testament to their Ops/Incident response teams and internal processes, it builds confidence in their ability to respond quickly when something does go wrong. Incredible work!
I feel like others lose opportunities by not doing the same. By publishing early and publishing the details they: keep the company in the news with positive stuff (free ad), get an internal documentation of the incident (ignoring the customer oriented "we're sorry" part), effectively get a free recruitment post (you're reading this because you're in tech and we do cool stuff, wink), release some internal architecture…
Additionally, these post-mittens work for Cloudflare because they have a great reputation and good uptime. If this were happening daily or weekly, it would be a warning sign to customers.
It’s a strategy other companies could adopt, but to do it effectively requires changes all across the organization.