Live data from Hacker News

Block Storage Issues Across All Regions: Incident Report for DigitalOcean

status.digitalocean.com

11–17 of 17 posts

Re: Block Storage Issues Across All Regions: Incident Report for DigitalOcean

#13

That's not seriously the real engineering portmortem is it? That looks more like a 'Resolved the issue' update - it is way too shallow and vague. If this was sent out at AWS as a COE (postmortem), it would be ripped apart - it is not going to satisfy anyone reading it that they should have confidence this class of failure isn't going to happen again. It looks like they haven't even identified the root cause(s) of the…

Yeah this is just a status update. Here is the TLDR:

> Efforts continue to find the incompatibility in the networking configuration change. Additionally, we are exploring improvements to our tools and processes to facilitate a finer grained, more incremental deployment method for wide, system-level changes.

Re: Block Storage Issues Across All Regions: Incident Report for DigitalOcean

#14
Adding my voice to the chorus here as a DO customer. This 188-word "postmortem" gives postmortems a bad name. I would like to know the details of the "network configuration change" and why it "caused incompatibilities". And also how you will ensure that this particular failure will not re-occur.

Trust and transparency are the currencies of the internet, in the same way that cigarettes and contraband are the currencies in prison. This post is worth approx. a 1/2 smoked cigarette.

Re: Block Storage Issues Across All Regions: Incident Report for DigitalOcean

#16
post #8

It's because of "reports" like these I didn't feel like staying as their customer. Whomever is in charge of [limiting scope and wording of] these reports should listen to a few things in private at their HQs.

Given that, would you trust the engineers from DO to work on your systems?

Of course, engineers are not individually responsible for large outages like that, I trust they always do their best (I'm a SRE myself).

Re: Block Storage Issues Across All Regions: Incident Report for DigitalOcean

#17
post #8

Earlier quoted context omitted.

Given that, would you trust the engineers from DO to work on your systems?

Of course, engineers are not individually responsible for large outages like that, I trust they always do their best (I'm a SRE myself).

Even if it’s the culture at DO to not question existing designs? I’m saying I would look at this example and question DO engineers that come to work for me rather than trust blindly that they won’t bring bad habits with them...
Post reply on HN