Live data from Hacker News

Today's Outage Post Mortem

blog.cloudflare.com

21–30 of 159 posts

Re: Today's Outage Post Mortem

#21
To the couldflare folks; It's refreshing to see you take responsibility, but I think you've been a bit too hard on yourselves by taking all the blame. First of all, what you hit was a unknown bug in JunOS, and Juniper is to blame for their part. Using some form of staging to slow roll-out of rule changes might have saved you from a full meltdown, but when you're getting attacked, every second counts. Slow versus fast roll-out is one of those really tough balancing acts in your situation. You did a great job with it; by the time I saw the "cloudflare is down" post in the newest queue, it was already back up running again.

Re: Today's Outage Post Mortem

#22
This is pretty impressive. Keep in mind most of the team is on the west coast so this happened at 1am on a Sunday and they put up a post mortem within hours. Obviously you would prefer it not happen at all, but that is a great response imo.

Re: Today's Outage Post Mortem

#23

So what are they going to change as a consequence? It seems logical to not rely on a single router vendor anymore, or to test new rules on a staging setup at least for a very short time before pushing them to all routers.

Adding different vendor types creates a cartesian explosion of possible combinations of bugs.

For standalone units that don't interact, n-version redundancy is good.

But if they have to interact and somebody has to troubleshoot, n-version redundancy is a nightmare.

The only reason it got fixed this quickly is because they had an intimate knowledge of a single vendor's products. That would be much more difficult with multiple vendors.

Re: Today's Outage Post Mortem

#24

This is pretty impressive. Keep in mind most of the team is on the west coast so this happened at 1am on a Sunday and they put up a post mortem within hours. Obviously you would prefer it not happen at all, but that is a great response imo.

Everyone who was needed (management, network operations, and the technical support team) were woken up as soon as this happened. The small team who monitor things during the night were quickly calling people. Some folks physically went to the office to help answers phones; others drove to one of our data centers.

Re: Today's Outage Post Mortem

#25
Impressive response. 30 minute outage for something most of the hosts I've worked with in the past would have been mystified about for hours. Then a quick RFO and promise of proactive SLA adjustments? Next time I need a CDN or attack mitigation I'll be talking to Cloudflare

Re: Today's Outage Post Mortem

#26
As always I'm glad to see Cloudflare post such detailed outage reports. They are one of the few providers I know of that is willing to go into such depth and that is one of the things I appreciate about them. That said, the outage that occurred was one that was indeed fully preventable. We don't exactly have as many locations as they do, but for internal resources at least, not pushing configuration changes to all devices (network included) is pretty standard practice. Basically I imagine for them a good routine to follow might be to script changes so that they are 'rolled out', something along the lines of push manual changes to a scripted 'random' router set (one in country A,B,C), wait 15 minutes and then push to the remaining router sets. That wouldn't work for all situations, such as if the entire network is seeing a DDoS or what have you, but I imagine they could adapt a routine that would prevent this particular scenario.

With all of that said, as a Cloudflare customer and also having a call with them tomorrow scheduled already over the WAF stuff, I find it a bit... frustrating that this is occurring now and such a kind of mistake.

Edit: As an aside, I wonder if the Puppet module for Junos will be extended to support route statements. That would make this kind of deployment much easier.

Re: Today's Outage Post Mortem

#27

So as far the DOS attack was very successful. It took down the site which it intended to and take down the network with lots and lots of the sites. Hope lessons are learnt and your next generation is less prone to these attacks

If the attacker knew about the Juniper bug and thought about a way to convince the network operators to introduce the rule of death themselves, then this is a nice hack indeed. It won't be easy for CloudFare to generically protect against these types of attacks. They could either have mechanisms to revert configurations faster or a way to test new configurations on a single router.

even according to their admission, if they had not made any change the apps would all have run, just possibly some lag, but making a change for malicious user without knowing the consequence lead to this scenario.

Re: Today's Outage Post Mortem

#28
> CloudFlare currently runs 23 data centers worldwide.

Shouldn't that always say - CloudFlare currently runs in 23 data centers worldwide?

Or is that just how one would phrase that if you rent multiple racks or a cage in a datacenter? ...because I've seen that a bunch of times before from just about everyone.

Just curious.

Re: Today's Outage Post Mortem

#30
post #21

To the couldflare folks; It's refreshing to see you take responsibility, but I think you've been a bit too hard on yourselves by taking all the blame. First of all, what you hit was a unknown bug in JunOS, and Juniper is to blame for their part. Using some form of staging to slow roll-out of rule changes might have saved you from a full meltdown, but when you're getting attacked, every second counts. Slow versus fast…

"by the time I saw the "cloudflare is down" post in the newest queue, it was already back up running again."

Not sure what your timeline shows but the "cloudflare is down" post hit the #1 spot on the front page just a few minutes after they went down. About 40 minutes after that, the services started to come back online for me.

That's a significant outage. That's not reflecting on the job they did bringing things back online but your statement made it seem like a minor outage.

Post reply on HN