Live data from Hacker News

DigitalOcean block storage is down

status.digitalocean.com

11–20 of 88 posts

Re: DigitalOcean block storage is down

#13

What unholy thing did they do that broke it across 12 different datacenters, good lord.

This does seem to indicate a notable lack of isolation for the blast radius between DO datacenters. Would be interesting to see the post mortem.

I get the feeling that whoever writes the post-mortem is going to have a bit of pressure to assure folks that there is isolation going forward.

Re: DigitalOcean block storage is down

#16

Earlier quoted context omitted.

This does seem to indicate a notable lack of isolation for the blast radius between DO datacenters. Would be interesting to see the post mortem.

I get the feeling that whoever writes the post-mortem is going to have a bit of pressure to assure folks that there is isolation going forward.

Whoever broke it is going to feel significant pressure to actually isolate things too.

Re: DigitalOcean block storage is down

#17
post #6

Earlier quoted context omitted.

Probably the old "one command ran on everything"

tmux is a dangerous tool in the wrong hands.

to be fair, it's dangerous even in the best hands. mistakes happen but business processes need to be in place to prevent catastrophes...

every time i see something like this, my inclination is to blame the CTO, not the engineer who pulled the trigger.

Re: DigitalOcean block storage is down

#18

Earlier quoted context omitted.

This does seem to indicate a notable lack of isolation for the blast radius between DO datacenters. Would be interesting to see the post mortem.

I get the feeling that whoever writes the post-mortem is going to have a bit of pressure to assure folks that there is isolation going forward.

That would be a bad sign that there’s something wrong with the culture. I would hope for a postmortem that identified flaws that genuinely needed to be fixed.

Re: DigitalOcean block storage is down

#19

This is really down for more than 2 hours!!!

Last night I was testing DO managed Kubernetes cluster with persistent volume claim and the volume took 15 minutes to reattach after the pod is rescheduled to another host. I thought it was just some weird hiccup and went to bed.

The incident report indicated the problem started 4 hours ago (around 9pm GMT) but I was having problem around 4pm. It's definitely not a 2-hour incident.

Re: DigitalOcean block storage is down

#20

Earlier quoted context omitted.

tmux is a dangerous tool in the wrong hands.

to be fair, it's dangerous even in the best hands. mistakes happen but business processes need to be in place to prevent catastrophes... every time i see something like this, my inclination is to blame the CTO, not the engineer who pulled the trigger.

A post mortem should always be a place to highlight deficiencies in processes and communicate necessary improvements put into place, not to blame. Blame should only occur if the cadence of outages becomes excessive. Complex systems are tricky, and to err is human.

Disclaimer: Ops/infra engineer in a previous life.

Post reply on HN