Live data from Hacker News

Gmail Services Global Outage

outage.report

101–110 of 175 posts

Re: Gmail Services Global Outage

#101
post #57

Earlier quoted context omitted.

Given how many services it affected and the mention of 404 errors, I would suspect a GFE bug or bad configuration that started rolling out worldwide (hence the geographically diverse, but not 100% spread of the issue). A decade ago, Tuesdays were a special day for GFEs, but that hasn't been the case for years now. Perhaps it's just the Tuesday curse that persists. :-) My money is either on that or the static content…

GFE?

Historically Google terminated all security at that front door layer- hope that changed by now. But interestingly a cert issue down the pipe could easily be the root cause.

Re: Gmail Services Global Outage

#102
post #58

We managed to centralize everything, email, git, even the web. I understand 99.99 looks fine, but is somewhat sad to see half the world without email.

I believe Google (G Suite) only claims 99.9% - sounds small, is actually a huge difference. They typically achieve a better % but that is all they "guarantee".

"Service Credit" means: (a) 10% of the total invoice charges for the affected month if the Monthly Uptime Percentage for any calendar month is between 99.0% and 99.9%; or (b) 25% of the total invoice charges for the affected month if the Monthly Uptime Percentage for any calendar month is between 99.0% and 95.0 %; or (c) 50% of the total invoice charges for the affected month if the Monthly Uptime Percentage for any calendar month is less than 95.0%.

Source: https://gsuite.google.com/terms/partner_sla.html

Edit: It appears Microsoft/Office 365 offers the same SLA / 99.9%

Re: Gmail Services Global Outage

#103

Earlier quoted context omitted.

Why not use italics instead? How does the ">" help? You can put it at the beginning of the paragraph, but lines get re-flowed so you can't put it at the beginning of lines.

The > was the original (?) convention for email. But seriously HN should just allow basic quotes. And, while they are at it, increase the size of voting arrows to make it possible to reliably hit them on mobile. I get it’s somewhat nice to have a site that does not constantly redesign everything, but pretending you lost the password for the server is just overdoing it in the opposite direction.

It took me two attempts to upvote this comment. :)

At least I didn't accidentally Flag something today.

Re: Gmail Services Global Outage

#104
post #87

Earlier quoted context omitted.

Not how it works. If you service has that 99.999% availability and you get a single unavailability event in a year, that's already 5 minutes of downtime completely independent from all other events that users may or may not experience, there is almost no chance of overlap between them. Users definitely notice that. And worse, events are going to be even less frequent than that and ever more noticeable and on top of t…

I think you have a misconception on what actual reliability is for more products and services. 99.999% is a solid service, 99.99999% is a hard to achieve target for enterprise software. To say Google is not the place to look for reliability is a pretty comical statement.

I don't have a misconception. Can you name a single internet service that has an actual five nines availability? That's definitely not google search nor gmail.

Re: Gmail Services Global Outage

#105
post #17

Earlier quoted context omitted.

From the first SRE book [1]: "The error budget stems from the observation that 100% is the wrong reliability target for basically everything (pacemakers and anti-lock brakes being notable exceptions). In general, for any software service or system, 100% is not the right reliability target because no user can tell the difference between a system being 100% available and 99.999% available. There are many other systems…

We don't tolerate houses collapsing out of nowhere, brakes failing over the course of normal usage and planes falling out of the sky during routine flights. But for some reason, we HAVE TO tolerate software crapping itself once a year? I don't accept this logic. This is just a sign of how sloppy the industry has become. This is the reason your phone becomes obsolete after 2 years, whereas your car can continue to run…

We actually do tolerate it. Plenty of critical parts in your car are designed to not be 100% available even in all expected cases.

For example, plenty of higher end cars in california come with summer tires that can't be used in cold weather/ice.

Even the brakes you are talking about must be replaced every X miles (depending on how new the car is, this may be between 10k and 50k miles)

Houses are not definitely designed to be 100% available. This is in fact why they fail due to fire or earthquake or other events. The design point is not instant failure, but it's also not "100% available".

Like the SRE book says, they make a tradeoff.

Re: Gmail Services Global Outage

#107

Earlier quoted context omitted.

I think you have a misconception on what actual reliability is for more products and services. 99.999% is a solid service, 99.99999% is a hard to achieve target for enterprise software. To say Google is not the place to look for reliability is a pretty comical statement.

I don't have a misconception. Can you name a single internet service that has an actual five nines availability? That's definitely not google search nor gmail.

I never claimed Google has 5 9s. Your claim that 99.999% for Gmail means Google isn't a place we should go to for reliability advice is comical at best, ignorant at worst.

Re: Gmail Services Global Outage

#108
post #35
post #27

Well Microsoft had their annual mail outage the other day... I guess they could have coordinated a bit better to get them done and out of the way at the same time.

Microsoft has had at least 2 major Azure outages that affected their SSO product. My system was unreachable by anyone but admins. Our systems engineer couldn’t do anything but wait on Microsoft to fix the issue on their end.

We spent ~4 hours troubleshooting an Office365 deployment before realizing the SSO outage was the cause.

The simultaneous fury and relief when we successfully logged users on the next day having made no functional changes was a watershed moment in our self-hosted services commitment.

Re: Gmail Services Global Outage

#109
post #58

We managed to centralize everything, email, git, even the web. I understand 99.99 looks fine, but is somewhat sad to see half the world without email.

I'd be happy to move away from Gmail. Unfortunatelty, I happen to hate spam so that's not really an option.

My Gmail account gets more spam in a week than I've gotten in the two years I've had FastMail. And Gmail's spam filter is surprisingly easy to defeat, considering. For months, putting "- -" at the start of the subject line would trick Gmail and put it in your normal inbox. (I just looked, and it looks like they've fixed it, but back in December it was bad.)

This myth that nobody but Gmail can handle spam needs to die. It was true in 2006, but it is not true in 2019.

Re: Gmail Services Global Outage

#110
post #54

Earlier quoted context omitted.

Please, please stop using mono space for quotes. It’s hard to read on small screens. Just > is better.

It’s not hard. It’s actually impossible. I agree, just use >. If I worked for YC as a HN mod I would literally spend a bit of time every day to review as many mono space-using comments I could every day and edit them to use > instead.

Or give us actual syntax for blockquotes? (> at the beginning of a line, markdown style, would be great…) Which I feel like gets used all the time in the discourse here, and for good reasons, too.

Then you would have that bit of time back, and the rest of us could stop scrolling back and forth when someone code-blocks a blockquote.

Post reply on HN