Today's Outage Post Mortem
blog.cloudflare.com
Today's Outage Post Mortem
1–10 of 159 posts
Re: Today's Outage Post Mortem
#2Re: Today's Outage Post Mortem
#3It appeals to my limited knowledge and non-existant experience that this would be a solution to the prevention of this occurring again in the future?
Re: Today's Outage Post Mortem
#4I'm not very educated on this end of the spectrum - but I wonder if a process is possible where a rule or router update of some description is applied to one router only, testing the specific schema before pushing to the rest of the routers, thereby failing one router and not failing the rest? I understand the need to respond as quickly as possible - but as stated in this case this was already a manual response. It a…
Re: Today's Outage Post Mortem
#5I'm not very educated on this end of the spectrum - but I wonder if a process is possible where a rule or router update of some description is applied to one router only, testing the specific schema before pushing to the rest of the routers, thereby failing one router and not failing the rest? I understand the need to respond as quickly as possible - but as stated in this case this was already a manual response. It a…
Re: Today's Outage Post Mortem
#6Re: Today's Outage Post Mortem
#7So what are they going to change as a consequence? It seems logical to not rely on a single router vendor anymore, or to test new rules on a staging setup at least for a very short time before pushing them to all routers.
It's probably more reasonable to split your network into a few more independent sections and never do updates which affect everything, but unless you're building the space shuttle (and can accept vastly higher costs and lower performance), it's probably better to pick one hardware platform, at least now.
Re: Today's Outage Post Mortem
#8Hope lessons are learnt and your next generation is less prone to these attacks
Re: Today's Outage Post Mortem
#9e.g., from simple things like hard drives in a raid from different vendors, to n-version programming in safety critical systems (like airplanes).
Re: Today's Outage Post Mortem
#10I'm not very educated on this end of the spectrum - but I wonder if a process is possible where a rule or router update of some description is applied to one router only, testing the specific schema before pushing to the rest of the routers, thereby failing one router and not failing the rest? I understand the need to respond as quickly as possible - but as stated in this case this was already a manual response. It a…
Don't sell yourself short: our ops team has been on our internal chat talking about how to do something exactly like this for the last hour or so. It's difficult at our scale to truly simulate traffic, but we should be able to roll rules out to just subsets of our network. That's already how we handle router OS upgrades. If a small handful of data centers had crashed, likely no one would have noticed because we've de…
My 2cents:
Any change should be considered dangerous, and be tested first, time weighted to it's level of change. (An internal policy that could be communicated publicly. One that I apply to all my staff)
It would be also good to have data center clusters (preferably a datacenter clusters are sharded among regions) which would allow these changes to happen as necessary. A random cluster being the "first cluster" with a roll back in place if fails, or proceed through to other clusters progressively until all are live.
The sharding should hopefully alleviate any corner of the world taking any massive hit due to degraded performance.