Block Storage Issues Across All Regions: Incident Report for DigitalOcean
status.digitalocean.com
Block Storage Issues Across All Regions: Incident Report for DigitalOcean
1–10 of 17 posts
Re: Block Storage Issues Across All Regions: Incident Report for DigitalOcean
#2> The outage was triggered as a result of a networking configuration change on the Block Storage clusters to improve handling packet loss scenarios. The new setting caused incompatibilities
So that doesn't tell us very much about the cause ("a networking configuration change") nor the effect ("incompatibilities").
Re: Block Storage Issues Across All Regions: Incident Report for DigitalOcean
#3I would have liked to hear more about how they are going to reduce the blast radius of such a change, because it sounds like something that could have been deployed to a single datacenter first
Re: Block Storage Issues Across All Regions: Incident Report for DigitalOcean
#4If this was sent out at AWS as a COE (postmortem), it would be ripped apart - it is not going to satisfy anyone reading it that they should have confidence this class of failure isn't going to happen again. It looks like they haven't even identified the root cause(s) of the failure...
Re: Block Storage Issues Across All Regions: Incident Report for DigitalOcean
#5Re: Block Storage Issues Across All Regions: Incident Report for DigitalOcean
#6Unfortunately, this doesn't explain why this change was applied to 5 datacenters at the same time, or why they didn't do that but it still affected 5 of them. I would have liked to hear more about how they are going to reduce the blast radius of such a change, because it sounds like something that could have been deployed to a single datacenter first
Re: Block Storage Issues Across All Regions: Incident Report for DigitalOcean
#7Re: Block Storage Issues Across All Regions: Incident Report for DigitalOcean
#8It's because of "reports" like these I didn't feel like staying as their customer. Whomever is in charge of [limiting scope and wording of] these reports should listen to a few things in private at their HQs.
Re: Block Storage Issues Across All Regions: Incident Report for DigitalOcean
#9Re: Block Storage Issues Across All Regions: Incident Report for DigitalOcean
#10I wonder if this just means someone changed an MTU configuration and it led to tons of fragmentation in different components of the network, and especially for any large file transfer making things timeout constantly to render an outage. Just a wild guess, but I’ve seen this happen with in-house datacenters before, so perhaps.