Live data from Hacker News

Google Outage in Europe

google.com

81–90 of 186 posts

Re: Google Outage in Europe

#81
post #14

Earlier quoted context omitted.

As time passes, the more big cloud providers there are, and the more complex they individually get (more products). Assuming the chance of an outage is actually static. If there is one (highly reliable and trusted) provider with five products one year, and three providers with ten products the next year, the chance of you seeing an outage has gone way up because of surface area.

Except that Google services don't run within GCP, but apart from it (mostly on Borg I'd guess).

If anyone else is curious, Borg is Google's cluster management software: https://kubernetes.io/blog/2015/04/borg-predecessor-to-kuber...

Re: Google Outage in Europe

#83
post #37

Earlier quoted context omitted.

Comments like this make me wonder if people really expect engineers to be fired because of an outage? I do not work at Google, but none of my workplaces would fire engineers because of a failure. Mistakes happen. As long as they are not repeated, everything is good. If your company fires people in situations like this, run away and never look back.

Googler here, not speaking on the behalf of the company, my opinions are my own People do absolutely NOT get fired over incidents. Making mistakes is human. An incident will prompt a review of the systems and safeguards in place to prevent such an incident, much like an airline incident investigation - basically "somebody fat-fingered it" is never the answer, postmortems are always blameless EDIT: now that I think of…

Yeah, this 'blameless' ethos has definitely trickled down from FAANG to decently-sized decently-reputed places I've worked at - and certainly to #EngTwitter.

I think it's a bit over-applied in some cases. Does it not commit you to the theorem that every process can be made so perfect as to be completely invulnerable to one human being making a mistake? (At least, in the form exemplified by the common tweets to the effect that "your processes are to blame for $incident, not your interns/engineers/etc".)

Even if you required two-person auth for every single thing, two people will make a mistake now and then, and in reality - due to our being social animals - the two probabilities are not truly independent.

I just don't see how this is feasible in reality. A more realistic principle feels like: "people will infrequently make mistakes, and that's of course natural and human and forgiveable, but far fewer incidents should be vulnerable to human error than currently are".

Re: Google Outage in Europe

#85
post #29
post #9

There were outages around the same time last year. Somebody in the HN thread commented back then that the employees evaluation and promotion window ends around december/eoy, thus more releases are made. https://en.m.wikipedia.org/wiki/Google_services_outages

Perf has been over for almost a month now, and the evaluation period was over more than two months ago.

> and the evaluation period was over more than two months ago

Excellent time to slack off a bit, make a few mistakes, then come next eval you can point to a marked improvement over the intervening six-to-nine months!

Re: Google Outage in Europe

#87
post #84

I do laugh at people that say "Using the cloud means less downtime than your own server", then things like this come along :D

Right?

My private website on my RPI is running now 2 Years without a problem and only minimal downtime due to rebooting for the new kernels.

It is amazing how much uptime you can achieve with a 5$ Computer in comparison to a 1730000000000$ (1,73 tera $) Company. Even if you compensate for dynamic content.

Re: Google Outage in Europe

#89
post #84

I do laugh at people that say "Using the cloud means less downtime than your own server", then things like this come along :D

It’s still less downtime than your own server and you have hundreds of engineers 24/7 ready to diagnose and fix issues

>It’s still less downtime than your own server

what makes you think so, actually?

Post reply on HN