Live data from Hacker News

Why the “Digital Ocean killed my company” incident scares the hell out of me

blog.checklyhq.com

151–160 of 185 posts

Re: Why the “Digital Ocean killed my company” incident scares the hell out of me

#151
post #81

I've had a DO mistake take down my site before, and when it was brought back up it had been reverted to several months prior. DO support was at a loss as to why this would have happened. I tried to restore from DO's backup service, but their backups had apparently stopped running several months prior as well. This was a major issue and could have easily been the death of my company, all because of a DO glitch. But it…

If you don't mind me asking, what is the preferred method to backup to S3? Is it possible to scp or rsync a mysqldump to S3 or do you install aws tools on DO and run aws s3 cp as a scheduled job?

I recommend looking into restic and/or rclone

Re: Why the “Digital Ocean killed my company” incident scares the hell out of me

#152
post #131

Earlier quoted context omitted.

Because hypothetically your instance should be hardware agnostic. If the physical hardware dies, it should automatically migrate to another physical server at their data center without your intervention. It will only look like an unexpected restart from your perspective. That's something worth paying for. It would be more comparable to two servers at Hetzner with rapid failover. But even that is more involved since y…

Right. I'm using stock Ubuntu LTS with Docker and my software is in docker images or otherwise easily deployable. Changes to the host system configuration are documented, so re-creating a host from scratch is not a problem. Things like automatic failover and high availability in general are cool, but your SLOs should be set according to your business model, e.g. they need to make business sense. I do understand how p…

I agree with you. Many applications grow linerally or not at all. I personally run a couple hundred servers, but it's all VMware, chef and clustered. Frankly, most just sit at a stupid low load average, but I don't fuss it because the spare cycles aren't wasted and I can have hosts and storage fail with impunity.

Many ways of doing it. Virtualization is basically all direct hardware access for the stuff that counts.

But I digress. I am guilty of over building things. In part for fun. In part hubris. In part I like sleeping at night

Re: Why the “Digital Ocean killed my company” incident scares the hell out of me

#153
post #145

I made this comment in an older thread but haven't gotten an answer on... the circumstances seem off to me: The thing that has me scratching my head is how this chain of events unfolded. I get that your fraud algorithm flagged it because of lack of established payment. how is this possible if what the tweet referred to as "locking us out of all of our backups and work"? surely an account history of any significance w…

Hi weaksauce... sorry I must have missed your earlier question. The account had been live for some time and in that sense had history but because of credits it didn't have payment history. As some others have commented lots of startups use credits to get their business going and depending on your usage they can last you for quite a while. Payment history indicates a willingness and capability to make payments. Part o…

I appreciate the response. a followup question to that would be how they got enough credits to be running for that long without any payment?

Re: Why the “Digital Ocean killed my company” incident scares the hell out of me

#154
post #32

I am old enough to have lived in soviet times for a short while. One thing I remember from back then - in order to get things done, you had to know someone. Perhaps your aunt's friend worked in the politbureau or your mom's classmate was a friend of the director. Through connections like that you could get what you needed. It was such a relief when times changed and regardless of who you were, you could start exchang…

Hey terryf - Zach here from DO. If you ever encounter an issue that you don't feel is getting the proper attention feel free to reach out. We constantly engage with developers on social, which is primarily Twitter, and we don't look at follower count to determine who to reply to. If for whatever reason you don't get the help you need, you can always email me directly (first name at).

Thanks, Zach

Re: Why the “Digital Ocean killed my company” incident scares the hell out of me

#155
post #79
post #36

Earlier quoted context omitted.

I'll toot my own theory and it is that nobody wants to pay for enough capable support staff, pay to keep the support staff at the ready often enough, pay support staff who want to stay in that role. They want to automate that all away as much as possible. The future is everyone who isn't somebody chatting about how they run their application on X... because they're mysteriously banned form Y and Z and the next guy ta…

>I'll toot my own theory and it is that nobody wants to pay for enough capable support staff, pay to keep the support staff at the ready often enough, pay support staff who want to stay in that role. The crazy thing is that it's really not that expensive. AWS, for example, offers 24x7 support w/ <1 hour response time for production issues starting at $100 a month. At worst it's 10% of your bill.

Hi treis - All DO customers have always received free, 24x7 Support. Of course, we're working on being as responsive as our customers deserve. A few months ago we also implemented a new paid Premier Support tier which features a live channel with 30-minute response times.

Here's some more information for anyone who's interested - https://www.digitalocean.com/support/#PremierSupport

Thanks, Zach, DigitialOcean Support

Re: Why the “Digital Ocean killed my company” incident scares the hell out of me

#156
post #36
post #32

I am old enough to have lived in soviet times for a short while. One thing I remember from back then - in order to get things done, you had to know someone. Perhaps your aunt's friend worked in the politbureau or your mom's classmate was a friend of the director. Through connections like that you could get what you needed. It was such a relief when times changed and regardless of who you were, you could start exchang…

I'll toot my own theory and it is that nobody wants to pay for enough capable support staff, pay to keep the support staff at the ready often enough, pay support staff who want to stay in that role. They want to automate that all away as much as possible. The future is everyone who isn't somebody chatting about how they run their application on X... because they're mysteriously banned form Y and Z and the next guy ta…

> nobody wants to pay for enough capable support staff, pay to keep the support staff at the ready often enough, pay support staff who want to stay in that role

It's not about money. I worked at a unicorn that paid very high wages for support staff AND allowed them to work remotely and asynchronously from anywhere in the world. If you lived in Southeast Asia or Eastern Europe you'd make more than a local doctor just answering emails.

We still simply couldn't hire enough halfway intelligent people fast enough to keep up with the user growth. For each support person we'd hire, there'd be 10,000 new customers joining the same week. "Automating that all away" was the only tractable way to respond to people at all in a reasonable time frame. Obviously the support quality was awful.

Re: Why the “Digital Ocean killed my company” incident scares the hell out of me

#157
post #115

Earlier quoted context omitted.

The CPU allocations on Lightsail are anemic compared to DO.

Are they? They come out fairly comparably here. I think Lightsail is a T3 instance "under the hood". https://joshtronic.com/2019/06/03/vps-showdown-digitalocean-...

Notice the memory and file I/O, huge gulfs in performance. The focus of my comment though was about Amazon's restrictive leaky bucket CPU allocation scheme, you can burn through your allocation of high cpu use and then get throttled to 5% of peak use iirc for hours while this bucket refills. DO lets you use much more CPU comparatively.

Re: Why the “Digital Ocean killed my company” incident scares the hell out of me

#158

Earlier quoted context omitted.

> Also, why would they want to? One of these days, they'll fuck up big time for many customers and get sued. They'll survive, but it'll cost a lot of money. Especially for Google (and other life-or-death services for many people), the solution seems kind of simple: charge for support. Google terminated your GMail-Account because you logged in from Turkey? Pay $50 to get somebody to listen to your story and work with…

> Pay $50 to get somebody to listen to your story and work with you on proving your identity. Personally, I would understand this as extortion.

Doesn't have to be $50, but a few might make sense to to keep their support from getting clogged with stuff that could be googled. Some sort of "Rescue Me" emergency flare option.

Re: Why the “Digital Ocean killed my company” incident scares the hell out of me

#159

Earlier quoted context omitted.

> Pay $50 to get somebody to listen to your story and work with you on proving your identity. Personally, I would understand this as extortion.

Doesn't have to be $50, but a few might make sense to to keep their support from getting clogged with stuff that could be googled. Some sort of "Rescue Me" emergency flare option.

It's billing for solving problems they have caused.

This is completely not Ok. If the idea is paid support for random problems, that's fine, if it's larger prices so it includes support, that's also fine. But if any company caused me a large damage and decided to ask money so they would reevaluate their actions, I'd go to the police.

Re: Why the “Digital Ocean killed my company” incident scares the hell out of me

#160

I've had a DO mistake take down my site before, and when it was brought back up it had been reverted to several months prior. DO support was at a loss as to why this would have happened. I tried to restore from DO's backup service, but their backups had apparently stopped running several months prior as well. This was a major issue and could have easily been the death of my company, all because of a DO glitch. But it…

> I don't see "I've never run a business before" as a valid excuse, nor do I see "the cloud vendor is better equipped to handle backups" as a valid excuse. If you're paying someone specifically to make backups for you, you should be able to trust that they've taken every reasonable measure to ensure that backups are actually being made and preserved.

You would think so, but then again you might be wrong. Better to be safe than sorry, no? I expected DO to make the backups I was paying them for and they didn't. Luckily I was making my own backups at the same time. Turns out my backups worked and theirs didn't. If I had just trusted them, I'd be out of business just like the dead company in question.

I'm not sure what's so confusing about "don't trust your vendors" but I've had to make this exact same reply way too many times.

Don't trust your vendors!

Post reply on HN