Live data from Hacker News

"DigitalOcean Killed Our Company"

twitter.com

361–370 of 620 posts

Re: "DigitalOcean Killed Our Company"

#361
post #307

Earlier quoted context omitted.

Mistakes happen, and algorithms are sometimes a necessary part of scale/efficiency. Everyone understands that. That said, what's highly troubling as a DO customer (and someone who is planning to deploy startup infrastructure of my own with DO) is: 1) The discrepancy between this customer's experience and clear assurances made on this very forum by high-level DO employees that: a. warnings are ALWAYS issued before sus…

Not sure why you’re being downvoted. Point 2 is very relevant. Scaling instances due to sudden peaks should be totally safe. Even when automated. Guess AWS is still lonely at the top.

AWS has default instance limits too, though, which you need to open ticket to increase.

Re: "DigitalOcean Killed Our Company"

#362

DigitalOcean is operating worse than a fly by night host (like AlphaRacks, GreenValueHost, etc). The reasonable course of action would've been to email the customer and throttle their API access to prevent load spikes, but DO instead locked their entire account (not just the service that DO felt was being abused). A fly by night will often only suspend the VM or database that is in question, not other services on the…

> worse than a fly by night host (like AlphaRacks

I saw a great deal on LEB for a KVM VPS from Alpharacks and signed up for a 2 year plan (my first mistake).

When SSHing in to the VPS it didn't have the advertised specs, and when I raised the issue with their support, they eventually fixed it..

Then I realized the second problem, they gave me the same IP address as someone else. You could still use the web VNC console, and as soon as you made an outbound network connection, inbound connections would work... for a few seconds... then SSH would drop. Reconnecting by SSH says "host key changed" i.e. you hit someone else's server sharing the same IP address. Using the web VNC console, works again for a few seconds, drops again.

It took about 7 days of arguing with their support to explain these two problems to them... by which time, the 3-day refund window had expired...

I admit some schadenfreude watching their recent disaster (all servers down since the last 2 weeks - https://www.lowendtalk.com/discussion/157613/popcorn-time-du... ).

Lesson learned about low-end boxes. I have had about 50/50 good and bad experiences, using half a dozen providers like this, not really making financial sense overall.

I recommend everyone stay far far away from AlphaRacks. If anything remains of them after this week.

Re: "DigitalOcean Killed Our Company"

#363
I think an important thing you can learn from this story is that you should keep your backup on a different host(s) or better even have replication enabled.

In these days, most apps generally can be migrated to a new host in seconds as long as you have the data source alive.

If they had access to thier data they probably should have been able to spin up a similar ec2 instance in minutes and say goodbye to DO forever.

Re: "DigitalOcean Killed Our Company"

#364
We had a similar incident with DO 8 days ago. It didn't kill our company, but we got hit hard.

Our business is Dynalist, an online outliner app. Many of our users store all their notes on Dynalist, so uptime is really important.

Starting 7 PM last Tuesday, we saw a slowdown in request handling. We filed a ticket with DO 2 hours after that (we also posted our initial tweet to keep our users informed: https://twitter.com/DynalistHQ/status/1131087411797270529).

A few hours later, we started to experience full downtime. Still no reply from DO. We filed another ticket with the prefix "[URGENT]". Still no reply.

We waited for 24 hours for their reply. We took turns taking naps because we're only a 2-person team.

After 24 hours, we tweeted @ DO (https://twitter.com/DynalistHQ/status/1131397013306847232). 2 hours later we finally got a support person working on our ticket. We didn't want to take it to the social media, but there doesn't seem to be any other way at that point. DO doesn't have phone support, and us "bumping" our support ticket didn't work either.

After 2 hours going back and forth on the support ticket and providing logs, DO's support person identified the issue and offered to move us to a less crowded server. They asked us what's a good time to do a manual migration if a live migration fails, and we replied immediately saying whenever is fine (we're experiencing downtime anyway).

We thought it's over, but we were so wrong.

They didn't reply in another 4 hours. That was 4 hours of more downtime. Sometimes, CPU steal is down a bit and our server could catch up some requests, although it would still take 10 seconds for our users to open Dynalist. But most of the time, our web app was totally inaccessible. Watching the charts on our dashboard go up and down felt like some of the hardest hours of my life... mainly because there's nothing we could do.

4 hours in, I realized we had to post another angry tweet to get a solution. There's nothing else to do other than trying to stay awake anyway. So I posted another tweet: https://twitter.com/DynalistHQ/status/1131497962184564737

This tweet didn't seem to work. Nothing happened in the next 3.5 hours and things started to feel surreal. I didn't know how much longer this downtime is going to last, and I didn't know what we were going to do about it.

At that time, it was 9:30 AM EDT and people were starting their day. We were getting more and more emails and tweets asking what is going on and where are their notes. A few customers were angry, but most were understanding and supportive.

At 9:55 AM EDT, DO finally did the live migration a few minutes before the time limit we gave them, which was 10 AM. That was the end of the incident; CPU steal was down to However, we couldn't trust DO any more. This weekend we're migrating to a dedicated server provider which has phone and live chat support. DO is pretty good for spinning up a $5 box quickly to test something, but we learned the hard way we shouldn't rely on it.

Our postmortem post: https://talk.dynalist.io/t/2019-05-22-dynalist-outage-post-m...

Re: "DigitalOcean Killed Our Company"

#365
post #270

As DigitalOcean's CTO, I'm very sorry for this situation and how it was handled. The account is now fully restored and we are doing an investigation of the incident. We are planning to post a public postmortem to provide full transparency for our customers and the community. This situation occurred due to false positives triggered by our internal fraud and abuse systems. While these situations are rare, they do happe…

Sure, but the email he received basically said "your account is locked. No other info. Thank You". That to me is a much scarier thing than anything else in the thread. How can anyone trust in your infrastructure if your standard protocol is literally just shutting down their entire operation without any form of review or communication?

You can't, obviously. Even though I've used them before I really doubt I'll ever use DigitalOcean again. I can almost understand terminating customers (with notice) via automated heuristics for suspicious behavior, especially on the low end of the hosting market, but locking out a legitimate paying customer from backups with no notice or recourse is terrifying.

Re: "DigitalOcean Killed Our Company"

#366
post #307

Earlier quoted context omitted.

Not sure why you’re being downvoted. Point 2 is very relevant. Scaling instances due to sudden peaks should be totally safe. Even when automated. Guess AWS is still lonely at the top.

AWS has default instance limits too, though, which you need to open ticket to increase.

Which is a much better policy than suspending the account.. 10 instances is nothing!

Re: "DigitalOcean Killed Our Company"

#367
post #273

Earlier quoted context omitted.

Old MSFT rule of thumbs was 2 bugs per day during bug crunch mode. Sounds crazy, but when you consider the number of "this text is wrong" and "that text box is too short" bugs that existed after a year of furious development, it wasn't too hard to achieve. Gotta hit that ZBB!

Brought back memories. I think it might be a little Stockholm syndrome but there was just something about the pressure of getting a release out when you know it only happens once every few years. Bug triage definitely improved my persuasion technique. Now its just "meh, we'll fix it in next months release".

Agile bug fix workflow ;)

Re: "DigitalOcean Killed Our Company"

#368
post #270

As DigitalOcean's CTO, I'm very sorry for this situation and how it was handled. The account is now fully restored and we are doing an investigation of the incident. We are planning to post a public postmortem to provide full transparency for our customers and the community. This situation occurred due to false positives triggered by our internal fraud and abuse systems. While these situations are rare, they do happe…

Thanks for the replies. Let me try to address a few of the things I have seen here. We haven't completed our investigation yet which will include details on the timeline, decisions made by our systems, our people, and our plans to address where we fell short. That said, I want to provide some information now rather than waiting for our full post-mortem analysis. A combination of factors, not just the usage patterns, led to the initial flag. We recognize and embrace our customers ability to spin up highly variable workloads, which would normally not lead to any issues. Clearly we messed up in this case.

Additionally, the steps taken in our response to the false positive did not follow our typical process. As part of our investigation, we are looking into our process and how we responded so we can improve upon this moving forward.

Re: "DigitalOcean Killed Our Company"

#369
You can not rely on one company. Let me repeat that. You. Can. Not. Rely. On. One. Company. If you are you’re negligent and you deserve what you get. You need your backups to be good and to be with a second provider. This is NOT rocket science. Treat your company seriously if you expect to sell services.

Re: "DigitalOcean Killed Our Company"

#370
post #322
post #201

Earlier quoted context omitted.

It's not just the volume of usage that can indicates fraud, but the pattern. In addition, relying on quotas creates a system that is easily games by perpetrators of fraud. Caps or quotas are not sufficient to deal with this problem.

If a cloud service is supposed to be ready for production, then customers should be safe to assume that they will not simply be shut down, especially not without warning. Otherwise, the provider must make clear that the service is only for hobby use and not for commercial use.

Every single major cloud service provider will shut you down without notice if they detect obvious fraud.
Post reply on HN