Live data from Hacker News

Gmail Services Global Outage

outage.report

131–140 of 175 posts

Re: Gmail Services Global Outage

#131

Earlier quoted context omitted.

I think you have a misconception on what actual reliability is for more products and services. 99.999% is a solid service, 99.99999% is a hard to achieve target for enterprise software. To say Google is not the place to look for reliability is a pretty comical statement.

I don't have a misconception. Can you name a single internet service that has an actual five nines availability? That's definitely not google search nor gmail.

Google cloud spanner has an availability SLA of 5 9s if you use multiple regions [0]. Not cheap though.

[0]: https://cloud.google.com/spanner/sla

Re: Gmail Services Global Outage

#132

Earlier quoted context omitted.

I believe Google (G Suite) only claims 99.9% - sounds small, is actually a huge difference. They typically achieve a better % but that is all they "guarantee". "Service Credit" means: (a) 10% of the total invoice charges for the affected month if the Monthly Uptime Percentage for any calendar month is between 99.0% and 99.9%; or (b) 25% of the total invoice charges for the affected month if the Monthly Uptime Percent…

For reference: over the course of a year, if you have an uptime of 99%, that means you are down for 3 days 15 hours per year. If you have 99.9% uptime, that means the downtime is roughly 9 hours per year. If you have 99.99% up, then the down is ~ 53 minutes per year. 99.999% then gives you ~5 minutes of down per year. 99.9999% will then give you ~31 seconds of downtime per year. One thing to keep in mind is that for…

> for each 9 of uptime you have, add about two Zeros to the budget

It's more like going from the stone age to the bronze age. You need expertise and architecture suitable for achieving high availability, these are your budget increases, they are not linear at all. But the infrastructure itself doesn't necessarily get more expensive, on the contrary, new architecture could allow you to use the cheapest stuff available on the market.

Re: Gmail Services Global Outage

#133
post #17

Earlier quoted context omitted.

From the first SRE book [1]: "The error budget stems from the observation that 100% is the wrong reliability target for basically everything (pacemakers and anti-lock brakes being notable exceptions). In general, for any software service or system, 100% is not the right reliability target because no user can tell the difference between a system being 100% available and 99.999% available. There are many other systems…

We don't tolerate houses collapsing out of nowhere, brakes failing over the course of normal usage and planes falling out of the sky during routine flights. But for some reason, we HAVE TO tolerate software crapping itself once a year? I don't accept this logic. This is just a sign of how sloppy the industry has become. This is the reason your phone becomes obsolete after 2 years, whereas your car can continue to run…

Consider that airplanes are relatively self-contained systems, whereas most of the systems we deal with in networking cross many different independent boundaries, each of which can independently fail for any number of reasons. There's more parties involved in regular operations of distributed software than in maintaining airplanes.

Re: Gmail Services Global Outage

#135

Earlier quoted context omitted.

I'd be happy to move away from Gmail. Unfortunatelty, I happen to hate spam so that's not really an option.

I’m on Fastmail. Spam is filtered out perfectly and no false positives in 2 years of service. Make the move away from Google, don’t be scared and don’t be guided by false excuses.

I hope you like paying per mailbox. That's a non-starter for someone like me who handles mail for 10-20 of my friends on my own domain.

That's $500-$1000 a year on Fastmail's Standard plan, the lowest tier plan that allows you to use your own domain.

Re: Gmail Services Global Outage

#136

Earlier quoted context omitted.

I think this is a false equivalency. If we're talking about "service unavailability", planes break all the time. Houses have to be vacated because of flooding, fire, insect infestation. Brakes do fail. Just like with software, we accept a certain level of risk in exchange for cost/convenience efficiencies (e.g. we don't want our planes to fall out of the sky, but we're okay with getting stranded in phoenix for 24 hou…

Also, brakes contribute to service unavailability. Brake pads need to be replaced on average every 50k miles, which takes the average driver 4 years. And let's say the average length of time your car is at the mechanic's to fix brakes is 3 days. That's 3 days of unavailability every 4 years just for brake pad replacements, or 99.8% availability (two nines!), just because of brake pad repairs. Add in all the other req…

Gmail seems to have 3 nines, although I couldn't find a better reference than this [1], where other services are included:

> [Google's infrastructure] delivers Gmail and other services to hundreds of millions of users with 99.978% availability and no scheduled downtime.

[1] https://support.google.com/googlecloud/answer/6056635?hl=en

PS. 99.978% availability translates as a downtime of ~ 2 hours/year total. Not bad! But it's when things break that we realize how performant and reliable they actually are.

Edits: various typos.

Re: Gmail Services Global Outage

#137

Earlier quoted context omitted.

I believe Google (G Suite) only claims 99.9% - sounds small, is actually a huge difference. They typically achieve a better % but that is all they "guarantee". "Service Credit" means: (a) 10% of the total invoice charges for the affected month if the Monthly Uptime Percentage for any calendar month is between 99.0% and 99.9%; or (b) 25% of the total invoice charges for the affected month if the Monthly Uptime Percent…

For reference: over the course of a year, if you have an uptime of 99%, that means you are down for 3 days 15 hours per year. If you have 99.9% uptime, that means the downtime is roughly 9 hours per year. If you have 99.99% up, then the down is ~ 53 minutes per year. 99.999% then gives you ~5 minutes of down per year. 99.9999% will then give you ~31 seconds of downtime per year. One thing to keep in mind is that for…

This is interesting! Is this a rule of thumb or common sense in the infrastructure world or is it something which there are studies? If there are studies do you have some links I could read further on?

Re: Gmail Services Global Outage

#138
post #58

We managed to centralize everything, email, git, even the web. I understand 99.99 looks fine, but is somewhat sad to see half the world without email.

I believe Google (G Suite) only claims 99.9% - sounds small, is actually a huge difference. They typically achieve a better % but that is all they "guarantee". "Service Credit" means: (a) 10% of the total invoice charges for the affected month if the Monthly Uptime Percentage for any calendar month is between 99.0% and 99.9%; or (b) 25% of the total invoice charges for the affected month if the Monthly Uptime Percent…

So even if it's down for the whole month I still have to pay (and probably can't even get out of it; even without some sort of longer contract the lock in is strong). Those are some big corporation terms all right.

Re: Gmail Services Global Outage

#139
post #127
post #58

We managed to centralize everything, email, git, even the web. I understand 99.99 looks fine, but is somewhat sad to see half the world without email.

Those of us who run their own mail servers: what is your MTBF? That is, ignoring all other differences, from the standpoint of pure availability, how do you compare to popular centralized email services? Just asking.

I had a serious fault two years ago, when several forest fires knocked out electric power for the two lines feeding the datacenter. Electricity failed three or four times in short periods, and for some reason this wrecked havoc on the UPS-generator combo, and power went out. We were down for six hours and mad at the co-location company.

Other than that, thirty minutes downtime 14 months ago when the public-facing redundant routers mis-detected a failover, failed STONITH and entered split-brain status.

I wouldn't commit to better than 99.9% with our infrastructure (single datacenter with remote data backups) but we exceed the metric most years (99.99% in all but one of the last ten years).

Re: Gmail Services Global Outage

#140
Ha. Apparently this site (outage.report) is used by scammers to lure victims.

>Discussion| Please don't call "support numbers" posted below — most probably it's a scam. Make sure to report and "downvote" such posts.

Post reply on HN