Live data from Hacker News

Google Cloud Networking Incident Postmortem

status.cloud.google.com

61–70 of 190 posts

Re: Google Cloud Networking Incident Postmortem

#61

I want a "24" style realtime movie of this event. Call it "Outage" and follow engineers across the globe struggling to bring back critical infrastructure.

it's pretty boring. real life computers aren't at all like hackers or csi:cyber. except for the skateboards, all real sysadmins ride skateboards.

Documentary that includes all the technical details wouldn't be. Kind of like the ambulance shows that seem to be popular now, but more technical.

Of course the target audience is probably tiny.

Re: Google Cloud Networking Incident Postmortem

#62

The outage lasted two days for our domain (edu, sw region). I understand that they are reporting a single day, 3-4 hours of serious issues but that’s not what we experienced. Great write up otherwise, glad they are sharing openly

What does your stack look like? It's hard to tailor a postmortem like this to everyone's individual experience but it is surprising to me that your experience is so different.

I know what you meant; however, reports should not be tailored to individual experience. The facts should be reported clearly. I’m happy they are open about the whole incident. -4 hours was more like two days for us.

Our stack? Multiple OC wan, 10G LAN with 1Gpbs clients. About 4,000+ users, EDU. We are super happy using Google. No complaints! Google is doing great.

Re: Google Cloud Networking Incident Postmortem

#64

Earlier quoted context omitted.

Unless I’m misunderstanding Google blog post they are reporting ~4+ hours of serious issues. We experienced about two days. If it was possible to have this fixed sooner I’m sure they would have done that. That’s not the point of my comment tough.

The root cause apparently lasted for ~4.5 hours, but residual effects were observed for days: > From Sunday 2 June, 2019 12:00 until Tuesday 4 June, 2019 11:30, 50% of service configuration push workflows failed ... Since Tuesday 4 June, 2019 11:30, service configuration pushes have been successful, but may take up to one hour to take effect. As a result, requests to new Endpoints services may return 500 errors for u…

That’s ok, I didn’t think your comment was dismissive. Those facts are buried in the report. Their opening sentence makes the incident sound lesser than what it really was.

Re: Google Cloud Networking Incident Postmortem

#65

Earlier quoted context omitted.

Is it real skateboards or boosted boards (or those one wheeled electric boards?).

I guess he mean the one true kind of sysadmins who's job contains moving physically in data center and deal with physical infrastructures. So it's real skateboard.

> So it's real skateboard.

Boosted boards are real skateboards too (https://boostedboards.com/) and would make moving through a DC even more effective ;)

Re: Google Cloud Networking Incident Postmortem

#66
post #13
post #2

Having only ever seen one major outage event in person (at a financial institution that hadn't yet come up with an incident response plan; cue three days of madness), I would love to be a fly on the wall at Google or other well-established engineering orgs when something like this goes down. I'd love to see the red binders come down off the shelf, people organize into incident response groups, and watch as a root cau…

It's interesting to see it go down. There's some chaos involved, but from my perspective it's the constructive[0] kind. If you're interested in how these sorts of incidents are managed, check out the SRE Book[1] - it has a chapter or two on this and many other related topics. Disclosure: I work in Google Cloud, but not SRE. [0]: https://principiadiscordia.com/book/70.php [1]: https://landing.google.com/sre/books/

Our own version of Netflix's "Chaos Money" is named "Eris" for precisely the reason mentioned in your first footnote.

Re: Google Cloud Networking Incident Postmortem

#67

Earlier quoted context omitted.

it's pretty boring. real life computers aren't at all like hackers or csi:cyber. except for the skateboards, all real sysadmins ride skateboards.

What?! It's the most exciting part of the job. Entire departments coming together, working as a team to problem solve under duress. What's more exciting than that?

I have done similar things several times and I think it would be boring.

It's Sunday so I guess they are not together. Instead there could be a lot of calls and working on some collaboration platforms. Everyone just staring at the screen, searching, reporting, testing and trying to shrink the problem scope.

If there's a record on everyone there must be a narrator explaining what's going on or audiences would definitely be confused.

It's Google so they have solid logging, analyzing and discovery means. Bad things do happen but they have the power to deal with them.

I suppose less technical firms(Equifax maybe?) encounter similar kind of crysis would be more fun to look at. Everything is a mess because they didn't build enough things to deal with them. And probably non-technical manager demanding precise response, or someone is blaming someone etc.

Re: Google Cloud Networking Incident Postmortem

#68
post #2

Having only ever seen one major outage event in person (at a financial institution that hadn't yet come up with an incident response plan; cue three days of madness), I would love to be a fly on the wall at Google or other well-established engineering orgs when something like this goes down. I'd love to see the red binders come down off the shelf, people organize into incident response groups, and watch as a root cau…

I used to be an SRE at Atlassian in Sydney on a team that regularly dealt with high-severity incidents, and I was an incident manager for probably 5-10 high severity Jira cloud incidents during my tenure too, so perhaps I can give some insight. I left because the SRE org in general at the time was too reactionary, but their incident response process was quite mature (perhaps by necessity).

The first thing I'll say is that most incident responses are reasonably uneventful and very procedural. You do some initial digging to figure out scope if it's not immediately obvious, make sure service owners have been paged, create incident communication channels (at least a slack room if not a physical war room) and you pull people into it. The majority of the time spent by the incident manager is on internal and external comms to stakeholders, making sure everyone is working on something (and often more importantly that nobody is working on something you don't know about), and generally making sure nobody is blocked.

To be honest, despite the fact that it's more often dealing with complex systems for which there is a higher rate of change and the failure modes are often surprising, the general sentiment in a well-run incident war room resembles black box recordings of pilots during emergencies. Cool, calm, and collected. Everyone in these kinds of orgs tend to quickly learn that panic doesn't help, so people tend to be pretty chill in my experience. I work in finance now in an org with no formally defined incident response process and the difference is pretty stark in the incidents I've been exposed to, generally more chaotic as you describe.

Re: Google Cloud Networking Incident Postmortem

#69

My burning question is what is a "relatively rare maintenance event type"?

I don’t have the inside knowledge of this outage but there are some details in here. They say that the job got descheduled due to misconfiguration. This implies the job could have been configured to serve through the maintenance event. It also implies there is a class of job which could not have done so. Power must have been at least mostly available, so it implies there was going to be some kind of rolling outage within the data center, which can be tolerated by certain workloads but not by others.

Re: Google Cloud Networking Incident Postmortem

#70
post #4

Why don't they refund every paid customer who was impacted? Why do they rely on the customer to self report the issue for a refund? For example GCS had 96% packet loss in us-west. So doesn't it make sense to refund every customer who had any API call to a GCS bucket on us-west during the outage?

Not directly GCS-related, but there was a big Youtube TV outage during the World Cup of last year (I think it was during semi-finals?). Google did apologize, but they only offered a free week of Youtube TV, which they implemented by charging me a week later than usual. I didn't feel compensated at all (it was a pretty important game that I missed!)
Post reply on HN