Live data from Hacker News

Cloudflare outage on June 21, 2022

blog.cloudflare.com

51–60 of 234 posts

Re: Cloudflare outage on June 21, 2022

#52

Yet another BGP caused outage. At some point we should collect all of them: - Cloudflare 2022 (this one) - Facebook 2021: https://news.ycombinator.com/item?id=28752131 - this one probably had the single biggest impact, since engineers got locked out of their systems, which made the fixing part look like a sci-fi movie - (Indirectly caused by BGP: Cloudflare 2020: https://blog.cloudflare.com/cloudflare-outage-on-july-…

Came here to say exactly this... things that mess with BGP have the power to wipe you off the internet.

Some more:

- Google 2016, configuration management bug/BGP: https://status.cloud.google.com/incident/compute/16007

- Valve 2015: https://www.thousandeyes.com/blog/steam-outage-monitor-data-...

- Cloudflare 2013: https://blog.cloudflare.com/todays-outage-post-mortem-82515/

Re: Cloudflare outage on June 21, 2022

#53
post #5

07:42: The last of the reverts has been completed. This was delayed as network engineers walked over each other's changes, reverting the previous reverts, causing the problem to re-appear sporadically. Ouch

This was something I was surprised not to see directly addressed in terms of follow up steps. When discussing process changes, they mention additional testing, but nothing to address what seems to be a significant communication gap.

Re: Cloudflare outage on June 21, 2022

#55

Really interesting that 19 cities handle 50% of the requests.

Well half of those cities were in Asia during business hours, so given that the majority of humans live in Asia it makes sense. CF data centers in Asia also seem to be less distributed than in the West (e.g. Vietnam traffic seems to go to Singapore) meanwhile CF has multiple centers distributed throughout the US.

Re: Cloudflare outage on June 21, 2022

#58

In a world where it can take weeks for other companies to publish a postmortem after an outage (if they ever do), I never ceases to amaze me how quickly CF manage to get something like this out. I think it's a testament to their Ops/Incident response teams and internal processes, it builds confidence in their ability to respond quickly when something does go wrong. Incredible work!

I feel like others lose opportunities by not doing the same. By publishing early and publishing the details they: keep the company in the news with positive stuff (free ad), get an internal documentation of the incident (ignoring the customer oriented "we're sorry" part), effectively get a free recruitment post (you're reading this because you're in tech and we do cool stuff, wink), release some internal architecture…

> I wonder how much those posts are calculated and how much organic/culture related.

Don't companies have a fiduciary duty to calculate things; the reason for doing something actually cannot just be that it's a nice thing to do? Not down to the word, but at least the general decision to be this way?

Re: Cloudflare outage on June 21, 2022

#60

In a world where it can take weeks for other companies to publish a postmortem after an outage (if they ever do), I never ceases to amaze me how quickly CF manage to get something like this out. I think it's a testament to their Ops/Incident response teams and internal processes, it builds confidence in their ability to respond quickly when something does go wrong. Incredible work!

I feel like others lose opportunities by not doing the same. By publishing early and publishing the details they: keep the company in the news with positive stuff (free ad), get an internal documentation of the incident (ignoring the customer oriented "we're sorry" part), effectively get a free recruitment post (you're reading this because you're in tech and we do cool stuff, wink), release some internal architecture…

I agree that this is a free ad/recruitment. However, it’s easy to see how more conservative businesses see this as a risk. They are highlighting their deficiencies, letting their big important clients know that human error can bring their network down.

Additionally, these post-mittens work for Cloudflare because they have a great reputation and good uptime. If this were happening daily or weekly, it would be a warning sign to customers.

It’s a strategy other companies could adopt, but to do it effectively requires changes all across the organization.

Post reply on HN