Live data from Hacker News

Cloudflare outage on June 21, 2022

blog.cloudflare.com

31–40 of 234 posts

Re: Cloudflare outage on June 21, 2022

#32
post #9

Am I the only who really doesn't think this is a big deal? They had an outage, they fixed it very quickly. Life goes on. Talking about the outage as if it's reason for us to all ditch CF, then buy/ run our own hardware (which will be totally better), so hyperbolic.

It was a bit of a thing as people in Europe started their office work, and found out a lot of their internet services were down, and they were unable to access the things they needed. It's rather dangerous that we all depend on this one service being online.

Re: Cloudflare outage on June 21, 2022

#33
post #5

07:42: The last of the reverts has been completed. This was delayed as network engineers walked over each other's changes, reverting the previous reverts, causing the problem to re-appear sporadically. Ouch

Well, the "we can't reach these data centers at all and need to go through the break glass procedure" was pretty "ouch" also.

Re: Cloudflare outage on June 21, 2022

#35

In a world where it can take weeks for other companies to publish a postmortem after an outage (if they ever do), I never ceases to amaze me how quickly CF manage to get something like this out. I think it's a testament to their Ops/Incident response teams and internal processes, it builds confidence in their ability to respond quickly when something does go wrong. Incredible work!

To further add to your point, the CTO is the one who shared it here & the CEO is incredibly active on forums & social media everywhere with customers. Communication has always been one of their strengths.

Re: Cloudflare outage on June 21, 2022

#36

Time and time again, this type of response proves that it's the right way handle a bad situation. Be humble, apologize, own your mistake, and give a transparent snapshot into what went wrong and how you're going to learn from the mistake. Or you could go the opposite direction and risk turning something like this into a PR death spiral.

Exactly. I trust businesses/people that are transparent about their mistakes/failures much more than the ones that avoid them (except Apple which never accepts their mistakes, but I still trust their products, I think I'm affected by RDF).

At the end of the day, everybody makes mistakes and that's okay. Everybody else also know that everybody makes mistakes. So why not accept it?

I really don't get what's wrong with accepting mistakes, learning from them, and moving on.

Re: Cloudflare outage on June 21, 2022

#40
Yet another BGP caused outage. At some point we should collect all of them:

- Cloudflare 2022 (this one)

- Facebook 2021: https://news.ycombinator.com/item?id=28752131 - this one probably had the single biggest impact, since engineers got locked out of their systems, which made the fixing part look like a sci-fi movie

- (Indirectly caused by BGP: Cloudflare 2020: https://blog.cloudflare.com/cloudflare-outage-on-july-17-202...)

- Google Cloud 2020: https://www.theregister.com/2020/12/16/google_europe_outage/

- IBM Cloud 2020: https://www.bleepingcomputer.com/news/technology/ibm-cloud-g...

- Cloudflare 2019: https://news.ycombinator.com/item?id=20262214

- Amazon 2018: https://www.techtarget.com/searchsecurity/news/252439945/BGP...

- AWS: https://www.thousandeyes.com/blog/route-leak-causes-amazon-a... (2015)

- Youtube: https://www.infoworld.com/article/2648947/youtube-outage-und... (2008)

And then there are incidents caused by hijacking: https://en.wikipedia.org/wiki/BGP_hijacking#:~:text=end%20us...

Post reply on HN