Live data from Hacker News

Google Cloud outage brings down Layer

status.layer.com

41–50 of 56 posts

Re: Google Cloud outage brings down Layer

#41
post #31

Earlier quoted context omitted.

Is it me, or are a lot of web-based service providers very chatty lately? I won't name and shame any particular ones, but I will say I've found myself regretting signing up for trials of certain services because of the almost sycophantic attention I'd receive from the oh-so-personable and friendly CEOs who make it a point to personally message all customers. I usually respond, initially, but then it quickly becomes p…

This happens because it works, though not necessarily so much for the HN crowd.

Does it though? In my experience a good product or service doesn't require constant spamming. It has nothing to with whether someone reads HN or not.

Re: Google Cloud outage brings down Layer

#42
post #2

Don't like the attitude. Pointing fingers doesn't help paying customers trapped by Layer poor design choices.

That was my first thought as well. Why did they need to start migrating customers to another AZ? I hope their customer's started asking that as well. The title should be "Poor Design Choices Brings Down Layer"

Re: Google Cloud outage brings down Layer

#44
post #31

Earlier quoted context omitted.

This happens because it works, though not necessarily so much for the HN crowd.

Does it though? In my experience a good product or service doesn't require constant spamming. It has nothing to with whether someone reads HN or not.

Does your request for confirmation incorporate the ancient HN discussion I linked?

Re: Google Cloud outage brings down Layer

#45
post #24

Earlier quoted context omitted.

You can't compare persistent disks failing in a whole zone, with a RAID array failing in a single machine. There is a reason why Amazon and Google takes EBS/Persistent Disk failures very seriously: there are not supposed to be unavailable during several hours, except if the whole datacenter is unable to operate (flood, fire, etc.), but it's not the case here. If your RAID fails, and you have a support contract which…

Two clarifications: the disks were not "unavailable", they had high latency (slow I/O) in one zone only (us-central1-a); and this affected only SSD PDs, not "regular" PDs. Per the SLA [1], it's "downtime" when PDs are completely unavailable for >5 minutes in at least two zones, and neither condition was met here. [1] https://cloud.google.com/compute/sla All that said, people choose SSD because it's faster and has hig…

this is a typical Google Cloud Support response (I used to host on GCloud). Stretching the definitions to somehow get out of responsibility. If the SSDs have super high latency, then for most purposes they are indeed 'unavailable'. There is a reason why the user provisioned SSDs and not a regular disk.

Re: Google Cloud outage brings down Layer

#47

Earlier quoted context omitted.

If you just go to layer.com, the first text you see on the page does a pretty good job of spelling out what it is. At least, it did for me. It's also more up-to-date than that comment, it would seem.

> Everything you need, from UI to infrastructure, to boost retention, engagement or drive transactions with the power of rich messaging. Wasn't enough for me. And if you click "Learn more" it's more marketing drivel. Granted my quora quote isn't much better.

Blake from Layer here. Have you taken a look at our developer documentation on developer.layer.com? I felt like we did a pretty good job of presenting the product capabilities. Our homepage and the developer documentation speak to different audiences. Let us know how the developer side matches up to your expectations.

Re: Google Cloud outage brings down Layer

#48

Most concise summary of Layer I could find on the internet quickly. > Layer is an amazingly elegant and light-weight solution for video communication. Layer is currently in a private beta primarily focused on Video, Voice and Chat on Android and iPhone. [1] Comment was in 2014. [1]: https://www.quora.com/What-is-the-difference-between-PubNub-...

Blake from Layer here: I'll definitely share your feedback with the product and marketing teams. Here's my version of a summary for a developer-centric audience:

Layer provides a comprehensive platform for adding rich messaging experiences inside other products. You can think of our offering as similar to iMessage or Facebook Messenger as a library / platform. We provide native SDKs on iOS and Android that provide a high-level development experience for implementing messaging. The SDK abstracts away all the low level details of implementing a great messaging system on mobile such as content synchronization and managing a persistent connection while still providing the developer with full control over the user experience. We also offer an open source UI toolkit called Atlas that provides a reference UI implementation on iOS and Android.

In addition to our mobile offering, we also provide a Javascript SDK for browser clients as well as raw REST and WebSocket APIs for other platforms. There is also a rich set of integration APIs in the form of backend to backend REST APIs and Webhooks for tracking events within the system.

The platform is fully managed and offered as a service. Historically our availability has been very strong, but last night exposed an achilles heal and we are working quickly to remediate the issues exposed.

Re: Google Cloud outage brings down Layer

#49
post #15

It doesn't do any good to point the finger at your vendors when your service goes down; that data isn't useful for your customers. Never forget the lesson of http://www.whoownsmyavailability.com/

Hey, Blake from Layer here: We were not at all trying to finger-point our issues at Google, only provide our customers with up to the minute transparency on where we were with the availability issue. We have received direct feedback from our customers that they do value detailed responses even when the news isn't great.

I take full responsibility for the issues here and the team is working to remediate the exposure as quickly as possible.

Re: Google Cloud outage brings down Layer

#50
post #4

> As we are now several hours into this outage and do not have satisfactory timeline for resolution, we have begun the process of migrating our hosts into another deployment zone within GCE. We will have a baseline set of services migrated within the hour and evaluate our ability to operate in a split deployment. Should we need to pursue a complete migration of hosts across zones then we would expect another 4-5 hour…

Hey, Blake from Layer here: we regularly undergo architecture and deployment reviews with our customers. We are fully transparent with the current deployment configuration and timelines for revisions.

Last night we lost a race to evolve our architecture and deployment ahead of a zone level issue that affected our total operations. We are working on it in earnest but there is very real complexity in operating a widely distributed real-time system.

Post reply on HN