Live data from Hacker News

Google Cloud outage brings down Layer

status.layer.com

21–30 of 56 posts

Re: Google Cloud outage brings down Layer

#21

I am so turned off when I click a "Pricing" link and get a contact form. Even more so when I read that, "our pricing team" [will get back to you]. So, you have an entire team of people who will try and maximize how much I pay? Sounds like a great experience doing business with you. /heavysarcasm

I automatically skip any product where I have to speak to a human at any time

Re: Google Cloud outage brings down Layer

#22
> As we are now several hours into this outage and do not have satisfactory timeline for resolution, we have begun the process of migrating our hosts into another deployment zone within GCE

Wait, what? Isn't running in multiple zones something like rule #1 or #3 in "how to run in the cloud"?

So why did they not already do this?

Re: Google Cloud outage brings down Layer

#23
post #17
post #15

It doesn't do any good to point the finger at your vendors when your service goes down; that data isn't useful for your customers. Never forget the lesson of http://www.whoownsmyavailability.com/

I'm not sure I agree. Customers like to know why it doesn't work. If it was a physical machine, they would have said something like "the disks are broken and we are replacing them". But it is cloud and they said "Google persistent disks are currently unavailable and they are fixing it".

But the real reason is "we didn't set up our system properly".

This is like saying "Hitachi Storage hard drives broke" when you actually mean "we didn't run RAID".

Re: Google Cloud outage brings down Layer

#24
post #23
post #17

Earlier quoted context omitted.

I'm not sure I agree. Customers like to know why it doesn't work. If it was a physical machine, they would have said something like "the disks are broken and we are replacing them". But it is cloud and they said "Google persistent disks are currently unavailable and they are fixing it".

But the real reason is "we didn't set up our system properly". This is like saying "Hitachi Storage hard drives broke" when you actually mean "we didn't run RAID".

You can't compare persistent disks failing in a whole zone, with a RAID array failing in a single machine.

There is a reason why Amazon and Google takes EBS/Persistent Disk failures very seriously: there are not supposed to be unavailable during several hours, except if the whole datacenter is unable to operate (flood, fire, etc.), but it's not the case here.

If your RAID fails, and you have a support contract which guarantees restoration within 1 hour, and it's not restored within 1 hour, then I think you can legitimately say something was wrong at your provider. It's not pointing fingers. Everyone does mistakes. It's taking responsibility.

That said, I agree they should have run in multiple zones, as recommended by Google, if they need/want to avoid that kind of downtime.

But I maintain Google Compute Engine Persistent Disk are not supposed to fail in such a way, and I'm quite sure Google will do whatever they can to avoid this in the future, instead of saying "don't point finger at us, it's supposed to happen".

Re: Google Cloud outage brings down Layer

#25
post #21

I am so turned off when I click a "Pricing" link and get a contact form. Even more so when I read that, "our pricing team" [will get back to you]. So, you have an entire team of people who will try and maximize how much I pay? Sounds like a great experience doing business with you. /heavysarcasm

I automatically skip any product where I have to speak to a human at any time

Is it me, or are a lot of web-based service providers very chatty lately?

I won't name and shame any particular ones, but I will say I've found myself regretting signing up for trials of certain services because of the almost sycophantic attention I'd receive from the oh-so-personable and friendly CEOs who make it a point to personally message all customers. I usually respond, initially, but then it quickly becomes pushy and intrusive, e.g. "Hi, I've noticed you haven't used [x] feature yet." "Hello? Are you getting my emails?" "Hello?"

I don't mean to be rude, but I didn't sign up for the "omg you're so friendly and amazingly helpful" show. I just wanted to try the service out. Kindly stop breathing down my neck! :/

Re: Google Cloud outage brings down Layer

#26
post #24
post #23

Earlier quoted context omitted.

But the real reason is "we didn't set up our system properly". This is like saying "Hitachi Storage hard drives broke" when you actually mean "we didn't run RAID".

You can't compare persistent disks failing in a whole zone, with a RAID array failing in a single machine. There is a reason why Amazon and Google takes EBS/Persistent Disk failures very seriously: there are not supposed to be unavailable during several hours, except if the whole datacenter is unable to operate (flood, fire, etc.), but it's not the case here. If your RAID fails, and you have a support contract which…

Two clarifications: the disks were not "unavailable", they had high latency (slow I/O) in one zone only (us-central1-a); and this affected only SSD PDs, not "regular" PDs. Per the SLA [1], it's "downtime" when PDs are completely unavailable for >5 minutes in at least two zones, and neither condition was met here.

[1] https://cloud.google.com/compute/sla

All that said, people choose SSD because it's faster and has higher throughput, so SSDs not being fast is obviously a real problem for applications relying on this, and rest assured we are indeed doing whatever we can to avoid this in the future.

Disclaimer: I work in Google Cloud Support.

Re: Google Cloud outage brings down Layer

#27
post #24
post #23

Earlier quoted context omitted.

But the real reason is "we didn't set up our system properly". This is like saying "Hitachi Storage hard drives broke" when you actually mean "we didn't run RAID".

You can't compare persistent disks failing in a whole zone, with a RAID array failing in a single machine. There is a reason why Amazon and Google takes EBS/Persistent Disk failures very seriously: there are not supposed to be unavailable during several hours, except if the whole datacenter is unable to operate (flood, fire, etc.), but it's not the case here. If your RAID fails, and you have a support contract which…

> That said, I agree they should have run in multiple zones, as recommended by Google

If you don't follow your vendor's recommendations for how to use their product, how can you blame them when that exact recommendation would have saved you?

> Google Compute Engine Persistent Disk are not supposed to fail in such a way, and I'm quite sure Google will do whatever they can to avoid this in the future

Sure. And the power to my office is not supposed to go out (and I've certainly worked in places where there has never been an unplanned power outage in decades), but if my business relies on it I need a UPS.

> instead of saying "don't point finger at us, it's supposed to happen".

It's not, and they shouldn't. Also unless you know something I don't, they didn't.

> If your RAID fails, and you have a support contract which guarantees restoration within 1 hour,

But as other commenter pointed out: Google did not violate the SLA during this, apparently. So…

Re: Google Cloud outage brings down Layer

#29

Earlier quoted context omitted.

If you just go to layer.com, the first text you see on the page does a pretty good job of spelling out what it is. At least, it did for me. It's also more up-to-date than that comment, it would seem.

Layer is just a building block for adding chat to your app. Similar to how you would use Elastic for search or Sendgrid for email.

That reminds me, I wonder what ever came of the Adria Richards v. Sendgrid issue.

Re: Google Cloud outage brings down Layer

#30
post #3

Earlier quoted context omitted.

Especially when they seem to be referencing only a single region. Multi-region deployments is the most basic protection against outages when using IaaS.

They should at least be in multiple availability zones. Multiple regions often comes with a lot of challenges, but there isn't much reason not to be redundant in multiple AZs.

> Multiple regions often comes with a lot of challenges

As an almost-customer of Layer (before their massive price increase), they led me to believe that this was one of the problems they would be solving for me. Nowhere on their website does it say, "We save money by not following best practices, so plan accordingly for occasional outages!"

Post reply on HN