Live data from Hacker News

Microsoft Azure suffers outage after cooling issue

datacenterdynamics.com

91–100 of 113 posts

Re: Microsoft Azure suffers outage after cooling issue

#91

Earlier quoted context omitted.

Google is fairly near MS (Azure did $1.9B in Q1, and GCP+Google Apps did $1.7B). But AWS is ~2.5x the other two.

Google Apps is not GCS. If you wanted to compare that number you'd have to throw in Office 365/Azure into the same number and it would dwarf Google. Google has ~3% market share compared to Microsoft's ~28% and Amazon's ~40%. Not even in the same league at the moment. Google is more on par with IBM and Rackspace, for now. Google will undoubtedly make strides in the space, but they haven't been tested.

> Google has ~3% market share compared to Microsoft's ~28% and Amazon's ~40%.

Where does this number come from?

If it is based on the revenue reported, be very careful with Microsoft's numbers. They report a lot of products as "Azure intelligent cloud", including Office suite subscriptions, on-premise server licences, and software (Windows, SQL Server) licensing revenue from other cloud providers in that number.

Pretty soon their claimed growth is going to flatten out, because they couldn't find any more revenue to report as "Azure intelligent cloud", like PC hardware ...

Re: Microsoft Azure suffers outage after cooling issue

#92
post #16
post #2

We're affected by this issue. And we had our alerts system in Azure as well, so we didn't get alerts about the outage (welp).

That's why T-Mobile's on-call engineers carry around AT&T phones. (source: friend who's an engineer at T-Mobile)

I remember that very well (worked for NSN), it wasn't funny. They learned that lesson from massive (unexpected) outage of their mobile network (2009), without possibility to reach technical persons: https://www.lightreading.com/ethernet-ip/server-glitch-crash...

Re: Microsoft Azure suffers outage after cooling issue

#93

VSTS is still down for us. TFS hosted code repos along with our entire bug system on VSTS means that no work is being done. I suspect we are gonna have to wait at least one other day at best for this to resolve. Meanwhile my local code goes even more out of sync. I’m probably just gonna spin up a git repo on my local machine and use that to share code with my team.

PM for VSTS here. The final scale unit in South Central US was brought back online a few hours ago, which means that the final accounts that were affected should now be operational. We're still restoring package management to some accounts, but otherwise, you should be back to working. Please feel free to reach out to me if you're still having trouble. Email is my HN username @ microsoft.com. We're _very_ sorry for this very significant outage.

Re: Microsoft Azure suffers outage after cooling issue

#94
post #52
post #48

Earlier quoted context omitted.

> Saying "check back in 2 hours" isn't useful. Having worked for a cloud provider, the reason they are saying that is because they are actively working to understand and fix the problem but haven't come to a well resound solution and thus they cannot give you a decent time estimate because you will probably get even more mad if they under/over estimate the time it took to fix it.

If they said this, "because they are actively working to understand and fix the problem but haven't come to a well resound solution and thus they cannot give you a decent time estimate because you will probably get even more mad if they under/over estimate the time it took to fix it." I would thank and applaud them. Tell me what it is you're doing at least. Why don't you understand the problem? What are you investiga…

I used to think like this too - e.g. I was happy when our national rail started announcing the cause of delays. But then a friend of mine was complaining that they did this, because he didn't want to be troubled with their internal problems - "just tell me what to do".

When your customer's demands are so directly opposed, you're somewhat caught between a rock and a hard place.

Re: Microsoft Azure suffers outage after cooling issue

#95

Earlier quoted context omitted.

I assume their build system checks for formatting and will raise an error if it doesn't conform. And this person would use the VS Code extension to auto-conform their code.

I'm questioning the soundness of such setup in general, and especially if it means that losing connection to a third-party prettifier makes you unable to work on your own codebase.

Hmm, the alternatives are not enforcing a similar code style, or enforcing it earlier on (e.g. on commit). I can understand why they would not want the former, and the latter is more annoying when experimenting, i.e. when code style does not matter that much yet. Thus, in CI sounds like the right choice.

Re: Microsoft Azure suffers outage after cooling issue

#96

So AWS has had some big outages, as has Azure. Has GCP had any big outages yet?

GCP has had multiple many-hour (6+) GLOBAL outages in the past year. I think it's at about 3 so far this year. But, it doesn't make the headlines like a 2-hour S3 outage in a single region, which must mean something ...

GCP's status history would seem to disagree, unless you have unusual definitions of "outage" and/or "global": https://status.cloud.google.com/summary

The last incident I'd personally classify as major lasted 39 minutes and was widely reported: https://status.cloud.google.com/incident/cloud-networking/18...

Disclaimer: I work at GCP but am not speaking for them. I also wish a speedy recovery for our colleagues at Azure: an outage like this can only result from many things going sideways simultaneously, and both the cause and recovery can be complicated in ways that flippant "well why didn't you just N+1 it" commenters here on HN can only guess at.

Re: Microsoft Azure suffers outage after cooling issue

#97

Visual Studio Online has been offline all day. They say it is due to the same Azure outage. This has had a productivity impact. If Microsoft didn't own GitHub, this may have prompted a move, but since they do it seems a little redundant given that Github will likely be on Azure too before long. https://blogs.msdn.microsoft.com/vsoservice/?p=17405

I think it's unlikely. Github were on cloud, and then moved to their own infrastructure a while back. In fact they have their own provisioning framework and all that fun stuff.

I doubt they will move back to Azure or any cloud for that matter. It's the same story with Dropbox and similar companies. Once past a point in scale, and depending on the case, for example need to control the data security to have certain certifications, it's essential to have your own infrastructure.

Re: Microsoft Azure suffers outage after cooling issue

#98
post #94
post #52

Earlier quoted context omitted.

If they said this, "because they are actively working to understand and fix the problem but haven't come to a well resound solution and thus they cannot give you a decent time estimate because you will probably get even more mad if they under/over estimate the time it took to fix it." I would thank and applaud them. Tell me what it is you're doing at least. Why don't you understand the problem? What are you investiga…

I used to think like this too - e.g. I was happy when our national rail started announcing the cause of delays. But then a friend of mine was complaining that they did this, because he didn't want to be troubled with their internal problems - "just tell me what to do". When your customer's demands are so directly opposed, you're somewhat caught between a rock and a hard place.

You can easily reflect both positions in your status page.

Those who don't need to be bothered with the details can refrain from reading them.

Re: Microsoft Azure suffers outage after cooling issue

#99
post #94

Earlier quoted context omitted.

I used to think like this too - e.g. I was happy when our national rail started announcing the cause of delays. But then a friend of mine was complaining that they did this, because he didn't want to be troubled with their internal problems - "just tell me what to do". When your customer's demands are so directly opposed, you're somewhat caught between a rock and a hard place.

You can easily reflect both positions in your status page. Those who don't need to be bothered with the details can refrain from reading them.

I really don't understand why everything has to be black or white with any of this stuff.

All this would take to keep both sides happy is a little link with "more info" below it.

Why things like this are so difficult, I will never understand.

Re: Microsoft Azure suffers outage after cooling issue

#100

Earlier quoted context omitted.

Just today I was having issues with the Prettier extension in VS Code, and I uninstalled it to see if that would fix it (I read that usually fixes the issues I was having). Then I realized that I couldn't install it again because VS Marketplace was down. This was like 8 hours ago and still no signs of recovery. Of course, all my builds are failing because of some stupid formatting issue that Prettier usually would so…

> Of course, all my builds are failing because of some stupid formatting issue that Prettier usually would solve, so yeah..thanks MSFT. Is failing builds due to formatting issues really a sound setup?

Yes.

When you have a style guide - test for it, if the test fails, then fail the style linter job and don't allow the change to be accepted.

It's failed to meet your acceptable code criteria after all.

If you find you are making your code unreadable just to pass, then your style guide is wrong. That needs fixing, not the CI job.

If you find an urgent "this needs to merge, style rules be dammed" change, allow your senior team members to overrule the style CI job and merge it anyway.

Post reply on HN