Live data from Hacker News

Microsoft Azure Outage

twitter.com

151–160 of 247 posts

Re: Microsoft Azure Outage

#151
post #125

> The issue is causing impact in waves, peaking approximately every 30 minutes. Does anyone have any general ideas on what kind of outage manifests itself like this? Devices retrying to authenticate every 30 minutes and finding the service is down perhaps?

Can sometimes be scaling/monitoring loops. i.e. cluster comes up, provides some limited service, gets overloaded and drops below required performance metric, gets killed by monitoring/scaling system, repeat...

Re: Microsoft Azure Outage

#152
post #137
post #41

Earlier quoted context omitted.

The point is PR. Never trust a status page if it's not directly connected to the monitoring system.

They never attach it to the monitoring because monitoring systems usually generate a lot of false positives which affect their published SLA.

Which impacts economics because some customers surely got deals guaranteeing some amount of credits based on up/downtime as reported by the status page.

And so updates to the status page become political and locked behind senior management approvals.. like AWS.

Re: Microsoft Azure Outage

#153
post #147

Earlier quoted context omitted.

Were you using Standard HDD disks? They have a really poor SLA, and are only usable for things like stateless VM Scale Sets or otherwise redundant services. We had to switch everything to SSD to get reliability comparable to on-prem VMware.

No entirely SSD. The problems stopped after a couple of weeks suddenly.

That sounds like what I've seen on Azure. Mystery weird problems we see, but they don't. Often in the network side. One time we were pretty sure they had a bad interface in a LAG group. Massive packet loss between hosts, but only on certain ephemeral source ports, about 1/8 of them.... Support couldn't find any issues even after a few days.

This was circa 2018 but AWS was so much more stable at that time. Ok, US-E-1 AWS had issues from time to time but they acked them and fixed them

Re: Microsoft Azure Outage

#155
post #90
post #34

This makes you wonder if some centralization patterns, i.e. Azure AD, are not a national security problem?

Azure AD is a nightmare. I don't know how many of you sign in to multiple tenants in the console, but it generally involves buying a new computer.

Firefox Multi-Account Containers extension. I couldn't live without it.

Re: Microsoft Azure Outage

#156
post #143

Earlier quoted context omitted.

For a development team, here's an example of something good about Azure: Microsoft gives us dev accounts with monthly Azure credits (e.g. $100) and you cannot spend more when those credits run out because there is no credit card etc. behind that account to charge the excess. Azure just like other cloud services (I've used AWS but as I understand it GCP is the same) doesn't believe in timely billing. You can and will…

> For a development team, here's an example of something good about Azure: Microsoft gives us dev accounts with monthly Azure credits (e.g. $100) First analogy I thought of were stories about drug dealers giving away free samples to schoolchildren to hook them up before asking for money.

Sure, it's obvious why they do this. Unlike drug dealers (who don't actually give school kids free crack, that makes no economic sense) it does make sense for Microsoft to ensure every kid who knows how to do rudimentary word processing knows Word, etc.

Nobody is under any illusion that Microsoft just really likes universities for some reason. But on the other hand, we did need lots of this stuff and it's very cheap, budgets are tight and it's not as though hand-rolling even more stuff would be cheaper - we do hand roll some things where it makes sense.

For example, periodically senior people say "Why do we spend $$$$ on a supercomputer? Surely we could rent one from the cloud?" and we (well, not me, different group same department) go OK, we will cost that for you. And they get Azure, Google, etc. to quote them for what they need a supercomputer to do, and then they present this, "The Cloud providers can do that for $$$$$". Ah, that's more money. No thanks, we will continue to run our own supercomputer.

It's not even close. Cloud supercomputer is great if you need the supercomputer for six weeks to do a special project and then you're done with it, the Cloud provider saves you a lot of money. But the University needs supercomputers all the time, so the numbers do not work.

Re: Microsoft Azure Outage

#157

Earlier quoted context omitted.

Azure Active Directory is not Active Directory but on Azure.

the people buying these things obviously have no idea about that. Migrating to Okta or something else neutral would cost the same, but hey, that's a different name

not even remotely close. okta for an enterprise is big dollars. most shops already have o365, so the AAD premium tier licensing is already paid for. aad and okta workforce are almost feature parity.

Re: Microsoft Azure Outage

#158
post #137
post #41

Earlier quoted context omitted.

The point is PR. Never trust a status page if it's not directly connected to the monitoring system.

They never attach it to the monitoring because monitoring systems usually generate a lot of false positives which affect their published SLA.

They never attach it to the monitoring because monitoring systems usually generate a lot of correct positives which affect their published SLA.

Works equally well. See the point?

Post reply on HN