Live data from Hacker News

Google Cloud outage brings down Layer

status.layer.com

51–56 of 56 posts

Re: Google Cloud outage brings down Layer

#51
post #7
post #4

> As we are now several hours into this outage and do not have satisfactory timeline for resolution, we have begun the process of migrating our hosts into another deployment zone within GCE. We will have a baseline set of services migrated within the hour and evaluate our ability to operate in a split deployment. Should we need to pursue a complete migration of hosts across zones then we would expect another 4-5 hour…

I imagine if one were a customer of this SaaS, it's on the customer to ask what availability to expect. Clearly this vendor thought that their savings on their IaaS bill outweighed any operational or reputational risk they'd suffer from an outage at a lower layer (pun unintended).

Blake from Layer here: We are forthright with all our customers about our current deployment configuration and the roadmap timelines for evolving into a deployment with higher availability characteristics. There is real complexity in operating a system such as ours in a widely distributed configuration and like any other company at our stage we regularly assess risks and make trade-offs. Sometimes we get things wrong.

We are very sorry to all our customers for the downstream impacts their businesses. We came up short and are doing everything we can to make it right.

Re: Google Cloud outage brings down Layer

#52
post #4

> As we are now several hours into this outage and do not have satisfactory timeline for resolution, we have begun the process of migrating our hosts into another deployment zone within GCE. We will have a baseline set of services migrated within the hour and evaluate our ability to operate in a split deployment. Should we need to pursue a complete migration of hosts across zones then we would expect another 4-5 hour…

I like how Algolia does that, https://www.algolia.com/infra (their blogposts and presentations go into much more detail) Currently thinking of creating a similar page for getstream.io, at the moment we always explain it during sales/onboarding calls. (we replicate our data to 3 different instances across multiple AZs)

Thanks for the link. We are taking a look at what Algolia has done here and will likely put together a public infrastructure overview page for Layer as well.

Re: Google Cloud outage brings down Layer

#53
post #2

Don't like the attitude. Pointing fingers doesn't help paying customers trapped by Layer poor design choices.

Blake from Layer here: I've reviewed the updates from last night and I don't feel like the tone was out of line. We were simply trying to provide our customers with complete transparency about where the issue was and where we were in restoring service.

With that said, we do feel that Google came up in short in their responses to us over the course of the issue. We pay handsomely on a support contract to get off-hours responses and issue escalations. The responses we received were hand-wavy and vague, leaving us without sufficient data to make decisions. We have raised these concerns with our Google representative and will be working with them to tighten our partnership going forward.

We take full responsibility for this event and are working to cover the exposure. Building a system and business with resource constraints and complex distributed technologies is a long game of managing risk and trade-offs. We're human and we make bad calls along the way. We are very sorry and violated our commitments to our customers and their users. The entire Layer engineering team is head down right now working to make it right.

Re: Google Cloud outage brings down Layer

#54

I am so turned off when I click a "Pricing" link and get a contact form. Even more so when I read that, "our pricing team" [will get back to you]. So, you have an entire team of people who will try and maximize how much I pay? Sounds like a great experience doing business with you. /heavysarcasm

I wonder if companies trying to make money should invest their time in talking to you.

Re: Google Cloud outage brings down Layer

#55
post #4

> As we are now several hours into this outage and do not have satisfactory timeline for resolution, we have begun the process of migrating our hosts into another deployment zone within GCE. We will have a baseline set of services migrated within the hour and evaluate our ability to operate in a split deployment. Should we need to pursue a complete migration of hosts across zones then we would expect another 4-5 hour…

I like how Algolia does that, https://www.algolia.com/infra (their blogposts and presentations go into much more detail) Currently thinking of creating a similar page for getstream.io, at the moment we always explain it during sales/onboarding calls. (we replicate our data to 3 different instances across multiple AZs)

More inspiration: https://www.mapbox.com/platform/
Post reply on HN