Live data from Hacker News

Cloudflare outage on June 21, 2022

blog.cloudflare.com

11–20 of 234 posts

Re: Cloudflare outage on June 21, 2022

#12

Are there any steps that can be taken to test these types of changes in a non-production environment?

It's very difficult if not impossible to create a staging environment that would well enough replicate production at this scale. What bog posts suggest as a remediation in the process: "There are several opportunities in our automation suite that would mitigate some or all of the impact seen from this event. Primarily, we will be concentrating on automation improvements that enforce an improved stagger policy for rollouts of network configuration and provide an automated “commit-confirm” rollback. The former enhancement would have significantly lessened the overall impact, and the latter would have greatly reduced the Time-to-Resolve during the incident."

Re: Cloudflare outage on June 21, 2022

#13
It's interesting that in 2022 we still have network issues caused by wrong order of rules.

Everybody at one time experiences the dreaded REJECT not being at the end of the rule stack but just too early.

Kudos to CF for such a good explanation of what caused the issue.

Re: Cloudflare outage on June 21, 2022

#14
I’m surprised they did not conclude roll outs should be executed over longer period with smaller batches. When a system is complicated as theirs with so much impact, the only sane strategy is slow rolling updates so that you can hit the brake when needed.

Re: Cloudflare outage on June 21, 2022

#17
post #9

Am I the only who really doesn't think this is a big deal? They had an outage, they fixed it very quickly. Life goes on. Talking about the outage as if it's reason for us to all ditch CF, then buy/ run our own hardware (which will be totally better), so hyperbolic.

> Talking about the outage as if it's reason for us to all ditch CF

at time of writing no comment has done that except you.

Re: Cloudflare outage on June 21, 2022

#18
post #14

I’m surprised they did not conclude roll outs should be executed over longer period with smaller batches. When a system is complicated as theirs with so much impact, the only sane strategy is slow rolling updates so that you can hit the brake when needed.

That's literally one of the conclusions.

Re: Cloudflare outage on June 21, 2022

#19

In a world where it can take weeks for other companies to publish a postmortem after an outage (if they ever do), I never ceases to amaze me how quickly CF manage to get something like this out. I think it's a testament to their Ops/Incident response teams and internal processes, it builds confidence in their ability to respond quickly when something does go wrong. Incredible work!

I feel like others lose opportunities by not doing the same. By publishing early and publishing the details they: keep the company in the news with positive stuff (free ad), get an internal documentation of the incident (ignoring the customer oriented "we're sorry" part), effectively get a free recruitment post (you're reading this because you're in tech and we do cool stuff, wink), release some internal architecture info that people will reference in discussions later. At a certain size it feels stupid not to post them publicly. I wonder how much those posts are calculated and how much organic/culture related.
Post reply on HN