Earlier quoted context omitted.
It could be good or bad; I suspect they've thought about it and have worked on succession (I hope!) and have like-minded people in the wings. But once it happens things will change and, to be honest, likely for the worse. edit> fix typo
> secession Succession?
Cloudflare outage on June 21, 2022
111–120 of 234 posts
Re: Cloudflare outage on June 21, 2022
#112In a world where it can take weeks for other companies to publish a postmortem after an outage (if they ever do), I never ceases to amaze me how quickly CF manage to get something like this out. I think it's a testament to their Ops/Incident response teams and internal processes, it builds confidence in their ability to respond quickly when something does go wrong. Incredible work!
I feel like others lose opportunities by not doing the same. By publishing early and publishing the details they: keep the company in the news with positive stuff (free ad), get an internal documentation of the incident (ignoring the customer oriented "we're sorry" part), effectively get a free recruitment post (you're reading this because you're in tech and we do cool stuff, wink), release some internal architecture…
I think their calculation (to the extent you can call it that) is that in the interest of PR and damage control, it’s better to get a thorough postmortem out quickly to stem the bleeding and keep people like us from going “I can’t wait to hear what happened at Cloudflare” for a week. Now we know, the customers have an explanation, and this bad news cycle has a higher chance of ending quickly.
Re: Cloudflare outage on June 21, 2022
#113Re: Cloudflare outage on June 21, 2022
#114Still seeing failed network calls. https://i.imgur.com/xHqvOzj.png
Is that actually coming from Cloudflare? iirc Cloudflare reports it self as Cloudflare not nginx in the 5xx error pages
Re: Cloudflare outage on June 21, 2022
#115Earlier quoted context omitted.
The internet runs on BGP, I would think that most internet issues would be a result of BGP then.
There are lots of other causes of incidents, like cut cables, failed router hardware, data centers losing power etc. It just seems that most of these are local enough and the Internet resilient enough that they don't cause global issues. Maybe the exception would be AWS us-east-1 outages :-)
Re: Cloudflare outage on June 21, 2022
#116Earlier quoted context omitted.
The internet runs on BGP, I would think that most internet issues would be a result of BGP then.
There are lots of other causes of incidents, like cut cables, failed router hardware, data centers losing power etc. It just seems that most of these are local enough and the Internet resilient enough that they don't cause global issues. Maybe the exception would be AWS us-east-1 outages :-)
Re: Cloudflare outage on June 21, 2022
#117Re: Cloudflare outage on June 21, 2022
#118BGP changes should be like the display resolution changes on your PC... It should revert as a failsafe if not confirmed within X minutes.
> Primarily, we will be concentrating on automation improvements ... and provide an automated “commit-confirm” rollback.
Re: Cloudflare outage on June 21, 2022
#119In a world where it can take weeks for other companies to publish a postmortem after an outage (if they ever do), I never ceases to amaze me how quickly CF manage to get something like this out. I think it's a testament to their Ops/Incident response teams and internal processes, it builds confidence in their ability to respond quickly when something does go wrong. Incredible work!
Re: Cloudflare outage on June 21, 2022
#120Yet another BGP caused outage. At some point we should collect all of them: - Cloudflare 2022 (this one) - Facebook 2021: https://news.ycombinator.com/item?id=28752131 - this one probably had the single biggest impact, since engineers got locked out of their systems, which made the fixing part look like a sci-fi movie - (Indirectly caused by BGP: Cloudflare 2020: https://blog.cloudflare.com/cloudflare-outage-on-july-…
> since engineers got locked out of their systems Sounds like the same happened here: "Due to this withdrawal, Cloudflare engineers experienced added difficulty in reaching the affected locations to revert the problematic change. We have backup procedures for handling such an event and used them to take control of the affected locations." But Cloudflare had sufficient backup connectivity to fix it. I'm curious how Cl…
Also lets face it - the utility of a trusted security guard/staff with an old fashioned physical key is pretty hard to screw up!