Live data from Hacker News

Google outage – resolved

news.ycombinator.com

761–770 of 870 posts

Re: Google outage – resolved

#761

Earlier quoted context omitted.

Same. At AWS, I once took an entire AZ down of a public-facing production service (with a mis-typed command), but that was nothing compared to when I accidentally deleted an entire region via internal console (too many browser tabs). Thank goodness turned out to be unused / unlaunched, non-production stack. I felt horrible for hours despite zero impact (in both the cases).

Irrelevant to the discussion, but I just wanted to say thank you for the categorized list of users I can follow on your profile!

Which tool do you use to follow users on HN?

Re: Google outage – resolved

#762
post #744

Earlier quoted context omitted.

Agreed, and also it's worth noting that we're talking about companies here. Yes, for any individual the amount of money lost is insane, but that's the risk for the company. If one individual can accidentally nearly bankrupt the company, then the company did not have proper risk management in place. That isn't too say that it wouldn't also affect my sleep quality.

Sadly a lot of managers don't see it this way, they'd rather assign blame.

They'd rather avoid the blame hitting them

Re: Google outage – resolved

#763

If you pay for Google Services, they have an SLA (service level agreement) of 99.9% [1]. If their services are down more than 43 minutes this month[2], you can request “free days” of usage. Edit: Services were down from ~12:55pm to ~1:52pm, it's 57minutes. Thanks hiby007 [1] https://workspace.google.com/intl/en/terms/sla.html [2] https://en.wikipedia.org/wiki/High_availability#Percentage_c...

If all 6 million G Suite customers, with an average number of users at 25 per G Suite account, paying the $20/user fee, requested the three day credit for this breach in the SLA contract for the outage, it'd cost Google about 300 million dollars.

Which is .22% of there COH this quarter...

Re: Google outage – resolved

#764

If you pay for Google Services, they have an SLA (service level agreement) of 99.9% [1]. If their services are down more than 43 minutes this month[2], you can request “free days” of usage. Edit: Services were down from ~12:55pm to ~1:52pm, it's 57minutes. Thanks hiby007 [1] https://workspace.google.com/intl/en/terms/sla.html [2] https://en.wikipedia.org/wiki/High_availability#Percentage_c...

If all 6 million G Suite customers, with an average number of users at 25 per G Suite account, paying the $20/user fee, requested the three day credit for this breach in the SLA contract for the outage, it'd cost Google about 300 million dollars. Which is .22% of there COH this quarter...

Or on the regular basis, they'd get $300M every other month (exclude any fee)

Re: Google outage – resolved

#765
post #743

Earlier quoted context omitted.

I've never seen an SLA that compensates for anything more than credit off your bill. I can't imagine a service that pays for loss of productivity, one outage and the whole company could be bankrupt. If your business depends on a cloud service for productivity you should have a backup plan if that service goes down.

I haven't seen one (at least for a SaaS company) that will compensate for loss of productivity/revenue etc, but something like Slack's SLA[0] seems like it's moving in the right direction. They guarantee a 99.99% uptime (max downtime of 4 min/22 seconds per month) and give 10x credits for any downtime. Granted, there's probably not many businesses that are losing major revenue because slack's down for half an hour, b…

It's hard to give a compensation for profit loss, as then you would have to know the profit of the customer beforehand and put an adequate pricing including that risk. It's almost like insurance!

Re: Google outage – resolved

#766

Given the blast radius of this (all regions appear to be impacted) along with the fact that services that don't rely on auth are working as normal, it must be a global authN/Z issue. I do not envy Google engineers right now.

> I do not envy Google engineers right now. A few years ago I released a bug in production that prevented users from logging into our desktop app. It affected about ~1k users before we found out and rolled back the release. I still remember a very cold feeling in my belly, barely could sleep that night. It is difficult to imagine what the people responsible for this are feeling right now.

Several years back when I was working at Google I made a mistake that caused some of the special results in the knowledge cards to become unclickable for a small subset of queries for about an hour. As part of the postmortem I had to calculate how many people likely tried to interact with it while it was broken. It was a lot and really made me realize the magnitude of an otherwise seemingly small production failure. My boss didn't give me a hard time, just pointed me toward documentation about how to write the report. And crunching the numbers is what really made me feel the weight of it. It was a good process.

I feel for the engineer who has to calculate the cost of this bug.

Re: Google outage – resolved

#767

Earlier quoted context omitted.

Oh man, you're right. Bloated gmail loaded instantly. What's going on? It's loading almost 2x to 3x faster.

Isn't this a good indication that the performance problem if gmail may not be related to the "bloat" of the frontend itself?

It might suggest that the frontend isn't the only issue, at least - and maybe this explains why it's usually so slow, if the frontend can be fast on a fast enough backend. On the other hand, the speed of the "basic HTML" version implies that the frontend can be the issue.

Re: Google outage – resolved

#769
post #665

Earlier quoted context omitted.

The big mistake in the system is that everyone in the world is relying on Google services... These problems would have less impact with a more diverse ecosystem.

Would they have less impact? Or would it have the same impact, just distributed across many more outages? You can rely on Google outages being very few and far between, and recovering pretty fast. For the benefits you get from such a connected ecosystem, I'm not sure anyone is net positive from using a variety of different tools rather than Google supplying many of them.

Compare closing down one road for repair a day per year to closing down all roads one day a year.

Re: Google outage – resolved

#770
post #654

4:41AM PT, Google services have been restored to my accounts (free & gsuite). And I have never seen them load so fast before - gmail progress bar barely seen for a fraction of a second whereas I am more used to seeing it for multiple seconds (2-3 sec) until it loads. I observe the same anecdotal speedup for other sites... drive, youtube, calendar. I wonder if they are throwing all the hardware they have at their serv…

Everything is snappier for a while if you turn it off and then on again

That's a great explanation for why I'm productive and so chirpy after a nap.
Post reply on HN