Microsoft Azure suffers outage after cooling issue
31–40 of 113 posts
Re: Microsoft Azure suffers outage after cooling issue
#32Earlier quoted context omitted.
That's why T-Mobile's on-call engineers carry around AT&T phones. (source: friend who's an engineer at T-Mobile)
That is a top 'did you know' factoid that I am sure I will tell others. But do AT+T engineers carry T-Mobile phones? If yes then they should put themselves together a deal so that none of the on-call engineers have to worry about running up big bills using their phones. When there are freak weather events they are all in it together.
Re: Microsoft Azure suffers outage after cooling issue
#33So AWS has had some big outages, as has Azure. Has GCP had any big outages yet?
AWS and Azure have had "big" outages people because actually use them. Rackspace and IBM are almost neck and neck with Google's best efforts (3% markshare Vs. 30%/40% for Azure/AWS)[0]. [0] https://www.skyhighnetworks.com/cloud-security-blog/microsof...
https://cloud.google.com/customers/
You know....services consumers actually use.
Re: Microsoft Azure suffers outage after cooling issue
#34Visual Studio Online has been offline all day. They say it is due to the same Azure outage. This has had a productivity impact. If Microsoft didn't own GitHub, this may have prompted a move, but since they do it seems a little redundant given that Github will likely be on Azure too before long. https://blogs.msdn.microsoft.com/vsoservice/?p=17405
Re: Microsoft Azure suffers outage after cooling issue
#35Re: Microsoft Azure suffers outage after cooling issue
#36Also unable to lodge a support ticket because the portal fails to identify me as having paid support (that API request appears to timeout).
Re: Microsoft Azure suffers outage after cooling issue
#37Earlier quoted context omitted.
They have some services that are "global", ie not tied to a given region. Those services' requests are actually processed all over the place, but south central is a big datacenter. The 9th biggest in the world, apparently. When it lost cooling and shut down, everything routed around it as planned... But it caused so much extra traffic that it overwhelmed the connections to other datacenters. The backlog of requests i…
> "Keep enough spare capacity around to handle losing one of the biggest datacenters in the world" is pretty unreasonable. Err what? It's entirely reasonable to expect Azure to handle the loss of a single DC and not have a 14+ hour global outage. I don't care how big the DC is, losing one should not take out the world, especially not for the length of time this one has been going on.
With that, though, it sounds like the size of this datacenter is way out of scale compared to the rest of their DC's. They are really going to need to break apart the services that they host there to make sure that DC to DC and region to region fail over works correctly.
Re: Microsoft Azure suffers outage after cooling issue
#38Earlier quoted context omitted.
They have some services that are "global", ie not tied to a given region. Those services' requests are actually processed all over the place, but south central is a big datacenter. The 9th biggest in the world, apparently. When it lost cooling and shut down, everything routed around it as planned... But it caused so much extra traffic that it overwhelmed the connections to other datacenters. The backlog of requests i…
> "Keep enough spare capacity around to handle losing one of the biggest datacenters in the world" is pretty unreasonable. Err what? It's entirely reasonable to expect Azure to handle the loss of a single DC and not have a 14+ hour global outage. I don't care how big the DC is, losing one should not take out the world, especially not for the length of time this one has been going on.
(I apologize if the following sounds snarky. I don't mean it that way, I just can't find better wording.)
Microsoft has repeatedly violated my sense of "reasonable" in the past, including in recent times with Windows 10. Therefore this kind of glitch isn't very shocking to me.