Live data from Hacker News

Cloudflare outage on June 21, 2022

blog.cloudflare.com

141–150 of 234 posts

Re: Cloudflare outage on June 21, 2022

#141

Yet another BGP caused outage. At some point we should collect all of them: - Cloudflare 2022 (this one) - Facebook 2021: https://news.ycombinator.com/item?id=28752131 - this one probably had the single biggest impact, since engineers got locked out of their systems, which made the fixing part look like a sci-fi movie - (Indirectly caused by BGP: Cloudflare 2020: https://blog.cloudflare.com/cloudflare-outage-on-july-…

Thats like blaming the hammer for breaking.

BGP is just a tool, it would be something else to do the same purpose.

Re: Cloudflare outage on June 21, 2022

#142
post #122

Earlier quoted context omitted.

The exception that proves the rule with Apple: https://appleinsider.com/articles/12/09/28/apple-ceo-tim-coo...

Yeah. Forgot that one. When it first came out it was terrible. Apparently so terrible that Apple apologized, perhaps for the first (and last) time for something.

They didn’t apologize about the direction the pro macs were going a few years back but they certainly listened and made amends for it with the recent Pro line and MacBook Pro enhancements

Re: Cloudflare outage on June 21, 2022

#144
post #141

Yet another BGP caused outage. At some point we should collect all of them: - Cloudflare 2022 (this one) - Facebook 2021: https://news.ycombinator.com/item?id=28752131 - this one probably had the single biggest impact, since engineers got locked out of their systems, which made the fixing part look like a sci-fi movie - (Indirectly caused by BGP: Cloudflare 2020: https://blog.cloudflare.com/cloudflare-outage-on-july-…

Thats like blaming the hammer for breaking. BGP is just a tool, it would be something else to do the same purpose.

Some tools are more fragile and error prone than others.

Re: Cloudflare outage on June 21, 2022

#145

In a world where it can take weeks for other companies to publish a postmortem after an outage (if they ever do), I never ceases to amaze me how quickly CF manage to get something like this out. I think it's a testament to their Ops/Incident response teams and internal processes, it builds confidence in their ability to respond quickly when something does go wrong. Incredible work!

I feel like others lose opportunities by not doing the same. By publishing early and publishing the details they: keep the company in the news with positive stuff (free ad), get an internal documentation of the incident (ignoring the customer oriented "we're sorry" part), effectively get a free recruitment post (you're reading this because you're in tech and we do cool stuff, wink), release some internal architecture…

>I feel like others lose opportunities by not doing the same

IMO it is a slippery slope to see this as opportunity too strongly. Sure, doing the right thing may be net beneficial to the business in the long run...but the $RIGHT_THING should be done first and foremost because it's the right thing.

Re: Cloudflare outage on June 21, 2022

#146

In a world where it can take weeks for other companies to publish a postmortem after an outage (if they ever do), I never ceases to amaze me how quickly CF manage to get something like this out. I think it's a testament to their Ops/Incident response teams and internal processes, it builds confidence in their ability to respond quickly when something does go wrong. Incredible work!

I'd love to see the postmortem from Facebook :(

Re: Cloudflare outage on June 21, 2022

#147

Earlier quoted context omitted.

That's the "commit-confirm" process they mention they will use in the write-up: > Primarily, we will be concentrating on automation improvements ... and provide an automated “commit-confirm” rollback.

Surprised everyone has not switched to this already - great idea

I assume there's some non-trivial caveats when using this with a widely-distributed system.

Re: Cloudflare outage on June 21, 2022

#148

In a world where it can take weeks for other companies to publish a postmortem after an outage (if they ever do), I never ceases to amaze me how quickly CF manage to get something like this out. I think it's a testament to their Ops/Incident response teams and internal processes, it builds confidence in their ability to respond quickly when something does go wrong. Incredible work!

To further add to your point, the CTO is the one who shared it here & the CEO is incredibly active on forums & social media everywhere with customers. Communication has always been one of their strengths.

To contrast this with the Atlassian outage recently is night and day.

Re: Cloudflare outage on June 21, 2022

#149
post #122

Earlier quoted context omitted.

The exception that proves the rule with Apple: https://appleinsider.com/articles/12/09/28/apple-ceo-tim-coo...

"Is it Apple Maps bad?" --Gavin Belson, Silicon Valley This one line will forever cement exactly how bad Apple Maps' release was. Thanks Mike Judge!

I agree, but lately (as in the past month) I've been finding myself using apple maps more and more than google. When on a complicated highway interchange, the 3d view that Apple Maps gives for which exit to take is a life-saver

Re: Cloudflare outage on June 21, 2022

#150
post #114

Earlier quoted context omitted.

Is that actually coming from Cloudflare? iirc Cloudflare reports it self as Cloudflare not nginx in the 5xx error pages

correct, i saw that too. the outage returned 500/nginx. no version number either on footer. @jgrahamc thought that was strange too as few commenters last night were caught off guard trying to determine if it was their systems or cloudflare. supposedly its been forwarded along.

yes, there is definitely an nginx service in the path. We don't have any nginx in our infrastructure, but this was the response we had for our urls during the outage.

500 Internal Server Error 500 Internal Server Error nginx

Post reply on HN