Today's Outage Post Mortem
21–30 of 159 posts
Re: Today's Outage Post Mortem
#22Re: Today's Outage Post Mortem
#23So what are they going to change as a consequence? It seems logical to not rely on a single router vendor anymore, or to test new rules on a staging setup at least for a very short time before pushing them to all routers.
For standalone units that don't interact, n-version redundancy is good.
But if they have to interact and somebody has to troubleshoot, n-version redundancy is a nightmare.
The only reason it got fixed this quickly is because they had an intimate knowledge of a single vendor's products. That would be much more difficult with multiple vendors.
Re: Today's Outage Post Mortem
#24This is pretty impressive. Keep in mind most of the team is on the west coast so this happened at 1am on a Sunday and they put up a post mortem within hours. Obviously you would prefer it not happen at all, but that is a great response imo.
Re: Today's Outage Post Mortem
#25Re: Today's Outage Post Mortem
#26With all of that said, as a Cloudflare customer and also having a call with them tomorrow scheduled already over the WAF stuff, I find it a bit... frustrating that this is occurring now and such a kind of mistake.
Edit: As an aside, I wonder if the Puppet module for Junos will be extended to support route statements. That would make this kind of deployment much easier.
Re: Today's Outage Post Mortem
#27So as far the DOS attack was very successful. It took down the site which it intended to and take down the network with lots and lots of the sites. Hope lessons are learnt and your next generation is less prone to these attacks
If the attacker knew about the Juniper bug and thought about a way to convince the network operators to introduce the rule of death themselves, then this is a nice hack indeed. It won't be easy for CloudFare to generically protect against these types of attacks. They could either have mechanisms to revert configurations faster or a way to test new configurations on a single router.
Re: Today's Outage Post Mortem
#28Shouldn't that always say - CloudFlare currently runs in 23 data centers worldwide?
Or is that just how one would phrase that if you rent multiple racks or a cage in a datacenter? ...because I've seen that a bunch of times before from just about everyone.
Just curious.
Re: Today's Outage Post Mortem
#29Re: Today's Outage Post Mortem
#30To the couldflare folks; It's refreshing to see you take responsibility, but I think you've been a bit too hard on yourselves by taking all the blame. First of all, what you hit was a unknown bug in JunOS, and Juniper is to blame for their part. Using some form of staging to slow roll-out of rule changes might have saved you from a full meltdown, but when you're getting attacked, every second counts. Slow versus fast…
Not sure what your timeline shows but the "cloudflare is down" post hit the #1 spot on the front page just a few minutes after they went down. About 40 minutes after that, the services started to come back online for me.
That's a significant outage. That's not reflecting on the job they did bringing things back online but your statement made it seem like a minor outage.