Live data from Hacker News

Microsoft Azure suffers outage after cooling issue

datacenterdynamics.com

41–50 of 113 posts

Re: Microsoft Azure suffers outage after cooling issue

#42
post #22

The Azure status page has more information. I suggest updating the link. > A severe weather event, including lightning strikes, occurred near one of the South Central US datacenters. This resulted in a power voltage increase that impacted cooling systems. Automated datacenter procedures to ensure data and hardware integrity went into effect and critical hardware entered a structured power down process. https://azure.…

I wonder if that's one of their facilities down here in San Antonio. Was getting flash flood alerts on my phone all night and morning.

Wow, I didn't tie these two events together until reading this comment. The flash floods last night were quite awful. mySA, a local news site (which I don't necessarily trust), has said that the daily rainfall total was 3x its historic record in the 1800s. [0]

It's always quite fascinating whenever cloud platforms like this have "leaky abstractions." GCP had a very long storage service degradation today, as well. [1] I don't know if it's related.

[0] https://www.mysanantonio.com/news/weather/article/Several-re...

[1] https://status.cloud.google.com/incident/storage/18003

edited for formatting

Re: Microsoft Azure suffers outage after cooling issue

#43

Edit: out my rant. It's been a long day because of this. Just going to leave it at that.

Just want to point out that once Azure launches their submersible datacenter units Azure Functions may literally become dead in the water.

One of the demos at Build was Azure Stack running on oil rigs, so you're not far off

Re: Microsoft Azure suffers outage after cooling issue

#44

Visual Studio Online has been offline all day. They say it is due to the same Azure outage. This has had a productivity impact. If Microsoft didn't own GitHub, this may have prompted a move, but since they do it seems a little redundant given that Github will likely be on Azure too before long. https://blogs.msdn.microsoft.com/vsoservice/?p=17405

[deleted]

Re: Microsoft Azure suffers outage after cooling issue

#45

So AWS has had some big outages, as has Azure. Has GCP had any big outages yet?

Outage is part of life, but Google's is most resilient in my experience.

Google doesn't have near the cloud presence of Amazon and Microsoft...maybe one day when they do, we can properly compare them. Given Google's small size/role in the space it's impossible to gauge if this is true.

Re: Microsoft Azure suffers outage after cooling issue

#48
post #40

The worst part has been the poor communication. If they were to give clearer insight from the get go, that'd give me more confidence and patience. Saying "check back in 2 hours" isn't useful.

> Saying "check back in 2 hours" isn't useful.

Having worked for a cloud provider, the reason they are saying that is because they are actively working to understand and fix the problem but haven't come to a well resound solution and thus they cannot give you a decent time estimate because you will probably get even more mad if they under/over estimate the time it took to fix it.

Re: Microsoft Azure suffers outage after cooling issue

#49
VSTS is still down for us. TFS hosted code repos along with our entire bug system on VSTS means that no work is being done.

I suspect we are gonna have to wait at least one other day at best for this to resolve. Meanwhile my local code goes even more out of sync.

I’m probably just gonna spin up a git repo on my local machine and use that to share code with my team.

Re: Microsoft Azure suffers outage after cooling issue

#50
post #48
post #40

The worst part has been the poor communication. If they were to give clearer insight from the get go, that'd give me more confidence and patience. Saying "check back in 2 hours" isn't useful.

> Saying "check back in 2 hours" isn't useful. Having worked for a cloud provider, the reason they are saying that is because they are actively working to understand and fix the problem but haven't come to a well resound solution and thus they cannot give you a decent time estimate because you will probably get even more mad if they under/over estimate the time it took to fix it.

Exactly, this happens even in just a normal production failure. I don't know what else they could have said/communicated. Not to mention this is the 7th largest data center in the world, resolving the problem likely took/is taking a long long time just because there are so many machines. I was lucky that the only outage effect I've suffered from is that my storage is locked, which means I can't add new file/edit code in production...but that's much better than it being down completely. My databases are geo-redundant, so that was a blessing today.
Post reply on HN