Live data from Hacker News

Cloudflare outage on June 21, 2022

blog.cloudflare.com

111–120 of 234 posts

Re: Cloudflare outage on June 21, 2022

#111
post #80

Earlier quoted context omitted.

It could be good or bad; I suspect they've thought about it and have worked on succession (I hope!) and have like-minded people in the wings. But once it happens things will change and, to be honest, likely for the worse. edit> fix typo

> secession Succession?

Eep yes, auto spell check on macOS is usually good but sometimes it causes a civil war.

Re: Cloudflare outage on June 21, 2022

#112

In a world where it can take weeks for other companies to publish a postmortem after an outage (if they ever do), I never ceases to amaze me how quickly CF manage to get something like this out. I think it's a testament to their Ops/Incident response teams and internal processes, it builds confidence in their ability to respond quickly when something does go wrong. Incredible work!

I feel like others lose opportunities by not doing the same. By publishing early and publishing the details they: keep the company in the news with positive stuff (free ad), get an internal documentation of the incident (ignoring the customer oriented "we're sorry" part), effectively get a free recruitment post (you're reading this because you're in tech and we do cool stuff, wink), release some internal architecture…

Ehhhh… I think it’s good (for us) that they do this, but I don’t think it’s a free ad (contrary to popular belief, not all news is good news, and this is bad news) and any sort of conversion rate on recruitment is probably vanishingly small (which would normally be fine, but incidents like these may turn off some actual customers, which is where actual revenue comes from).

I think their calculation (to the extent you can call it that) is that in the interest of PR and damage control, it’s better to get a thorough postmortem out quickly to stem the bleeding and keep people like us from going “I can’t wait to hear what happened at Cloudflare” for a week. Now we know, the customers have an explanation, and this bad news cycle has a higher chance of ending quickly.

Re: Cloudflare outage on June 21, 2022

#114

Still seeing failed network calls. https://i.imgur.com/xHqvOzj.png

Is that actually coming from Cloudflare? iirc Cloudflare reports it self as Cloudflare not nginx in the 5xx error pages

correct, i saw that too. the outage returned 500/nginx. no version number either on footer. @jgrahamc thought that was strange too as few commenters last night were caught off guard trying to determine if it was their systems or cloudflare. supposedly its been forwarded along.

Re: Cloudflare outage on June 21, 2022

#115
post #68

Earlier quoted context omitted.

The internet runs on BGP, I would think that most internet issues would be a result of BGP then.

There are lots of other causes of incidents, like cut cables, failed router hardware, data centers losing power etc. It just seems that most of these are local enough and the Internet resilient enough that they don't cause global issues. Maybe the exception would be AWS us-east-1 outages :-)

Maybe a testament to BGP's effectiveness that so many large-scale outages are due to misconfiguring BGP rather than the frequent cable cuts and hardware failures that BGP routes around.

Re: Cloudflare outage on June 21, 2022

#116
post #68

Earlier quoted context omitted.

The internet runs on BGP, I would think that most internet issues would be a result of BGP then.

There are lots of other causes of incidents, like cut cables, failed router hardware, data centers losing power etc. It just seems that most of these are local enough and the Internet resilient enough that they don't cause global issues. Maybe the exception would be AWS us-east-1 outages :-)

BGP is the reason you don't hear about cable cuts taking down the internet.

Re: Cloudflare outage on June 21, 2022

#118

BGP changes should be like the display resolution changes on your PC... It should revert as a failsafe if not confirmed within X minutes.

That's the "commit-confirm" process they mention they will use in the write-up:

> Primarily, we will be concentrating on automation improvements ... and provide an automated “commit-confirm” rollback.

Re: Cloudflare outage on June 21, 2022

#119

In a world where it can take weeks for other companies to publish a postmortem after an outage (if they ever do), I never ceases to amaze me how quickly CF manage to get something like this out. I think it's a testament to their Ops/Incident response teams and internal processes, it builds confidence in their ability to respond quickly when something does go wrong. Incredible work!

To be fair though they sort of MUST do things like this to have our confidence - their whole business is about being FAST and AVAILABLE. Were not talking about Oracle here :-D

Re: Cloudflare outage on June 21, 2022

#120
post #103

Yet another BGP caused outage. At some point we should collect all of them: - Cloudflare 2022 (this one) - Facebook 2021: https://news.ycombinator.com/item?id=28752131 - this one probably had the single biggest impact, since engineers got locked out of their systems, which made the fixing part look like a sci-fi movie - (Indirectly caused by BGP: Cloudflare 2020: https://blog.cloudflare.com/cloudflare-outage-on-july-…

> since engineers got locked out of their systems Sounds like the same happened here: "Due to this withdrawal, Cloudflare engineers experienced added difficulty in reaching the affected locations to revert the problematic change. We have backup procedures for handling such an event and used them to take control of the affected locations." But Cloudflare had sufficient backup connectivity to fix it. I'm curious how Cl…

Worst case if I was designing this I would probably have a satellite connection running over Iridium at each of their biggest DC's

Also lets face it - the utility of a trusted security guard/staff with an old fashioned physical key is pretty hard to screw up!

Post reply on HN