Live data from Hacker News

Microsoft Azure Outage

twitter.com

171–180 of 247 posts

Re: Microsoft Azure Outage

#171

Earlier quoted context omitted.

> I have zero sympathy for people being affected by this because if you voluntarily use Azure then this is what you deserve I guess many developers do not use Azure voluntarily but are forced to by their companies (or customers).

And the companies are forced to due to huge contracts with Microsoft

Literally the whole reason my last org got into Azure.

Lots of MSSQL and PowerBI licenses, lots of other Windows env features. Great deals to bundle those in w/ Azure deployments.

Great pricing too -- for the first 3 years. But at 4 years...

Re: Microsoft Azure Outage

#172
post #147

Earlier quoted context omitted.

No entirely SSD. The problems stopped after a couple of weeks suddenly.

That sounds like what I've seen on Azure. Mystery weird problems we see, but they don't. Often in the network side. One time we were pretty sure they had a bad interface in a LAG group. Massive packet loss between hosts, but only on certain ephemeral source ports, about 1/8 of them.... Support couldn't find any issues even after a few days. This was circa 2018 but AWS was so much more stable at that time. Ok, US-E-1…

Yes the lack of them being able to see any problems was a constant problem.

Our AWS reps are all over stuff when it goes down. I regularly get to talk to actual real product managers and engineers via our enterprise support if anything goes wrong.

Re: Microsoft Azure Outage

#173

What's the point of having a status page if it doesn't indicate the issues? https://status.azure.com/en-us/status Azure, Teams, Outlook are almost down from Greece and Germany, and their status page shows that everything is fine :-)

We have been here before...HN is the only status page that matters.

Re: Microsoft Azure Outage

#174

What's the point of having a status page if it doesn't indicate the issues? https://status.azure.com/en-us/status Azure, Teams, Outlook are almost down from Greece and Germany, and their status page shows that everything is fine :-)

When anything on that page turns not-green, there are news stories about it. Not positive ones. So exec approval is needed, because the decision to flip something on that page is ultimately the decision to cause stories negative to MS to be published. The exec has to weigh whether pissing off the customers (by failing to acknowledge reality) is worth the bad press and SLA fallout.

It has nothing to do with press. This is negative press already, and journalist can use this to write their stories without waiting for the official light to go from green to yellow.

It's about contractual obligations and SLAs. Things are not officially down in most agreements until MSFT acknowledges they're down. Refunds issued because your blob storage failed to meet 99.9999 uptime to your largest customers are directly tied to these statuses.

Re: Microsoft Azure Outage

#175
post #137
post #41

Earlier quoted context omitted.

The point is PR. Never trust a status page if it's not directly connected to the monitoring system.

They never attach it to the monitoring because monitoring systems usually generate a lot of false positives which affect their published SLA.

Which means if one were to require monitoring and status pages to be connected, one of two things happen (for each monitored component):

(1) The monitoring system would be altered to ignore tests that return false positives (at the expense of missing the alert when it represents an outage).

(2) Fixing the monitoring. It wasn't working for the sysadmins/operators, anyway, since it had so many false positives that their "mental model" was essentially based on (1), anyway.

At least, where I've forced the issue of doing just this, that's exactly what happened. At the end of the day, especially since SLAs took a hit and that affected bonus payouts, monitoring got a lot better -- as did overall team function when we truly realized how bad things were -- we stopped doing workarounds and started fixing problems at a more fundamental level which led to SLAs that were both accurate and excellent.

It helped bring attention to a hidden problem which resulted in time being allocated to fix tests that dropped constant false-positives and to evaluate each for whether or not it should exist in the first place.

Re: Microsoft Azure Outage

#176
post #90

Earlier quoted context omitted.

Azure AD is a nightmare. I don't know how many of you sign in to multiple tenants in the console, but it generally involves buying a new computer.

YES. I have the dubious honor of needing to use at least 4 different Teams tenants over the course of a week and it is enough to make me want to pitch my computer into the sea. App, browser, private browser - doesn't seem to matter. When I try to sign in, Microsoft will pick one of the tenants seemingly at random, regardless of what URL I use, and try to sign me in - of course, since there is usually no visual cue as…

Use browser profiles, choose a different profile picture for each, then use one profile per tenant. Done.

Re: Microsoft Azure Outage

#177
post #132

Earlier quoted context omitted.

If they were relying on Outlook and Teams to be productive, they probably couldn't really work before either.

What a naive comment. As if the only truly important jobs exist in engineering and require nothing but git and a book on C.

I interpreted this comment as more of a jab at how inefficient are outlook and teams themselves as applications.

I don't know if it's the right interpretation to have, but I kind of agreed with it, considering huge issues I had with teams (curiously some of them are only there for linux users, weird when considering the fact that I only use teams' web page) - not saying I could do better though!

Re: Microsoft Azure Outage

#178

Earlier quoted context omitted.

When anything on that page turns not-green, there are news stories about it. Not positive ones. So exec approval is needed, because the decision to flip something on that page is ultimately the decision to cause stories negative to MS to be published. The exec has to weigh whether pissing off the customers (by failing to acknowledge reality) is worth the bad press and SLA fallout.

Which means it's not a status page any more. Defeating the supposed purpose.

"SLA refund page"?

Re: Microsoft Azure Outage

#179
post #82

Earlier quoted context omitted.

> I have zero sympathy for people being affected by this because if you voluntarily use Azure then this is what you deserve I guess many developers do not use Azure voluntarily but are forced to by their companies (or customers).

We're migrating on Teams because of that kind of reasoning. It's utter shit of a service. Even worse if you need to write integrations for it

What will you use instead? What don't you like about teams (or wiring integrations)

Re: Microsoft Azure Outage

#180

Azure is the most developer hostile cloud environment. I have zero sympathy for people being affected by this because if you voluntarily use Azure then this is what you deserve. Sorry for being so miserable, but Azure has given me soooo much grief over the last 10 years that I'm just completely done with this shitshow of a platform.

I know Azure generally sucks... If you think you cannot go lower, you should try Oracle Cloud. That is a total piece of dung of a Cloud Service.

I tried it a couple of years ago. After finishing the trial, I removed all instances and disks, supposedly completely blanking the account. And also supposedly deleted the account.

To this day, I still keep receiving some kind of invoice for about $2 USD that they say I owe. And when I login into the "oracle cloud account" nothing works because my account seems to be half-deleted. (like I get error screens when accessing several of their piece of shit panels).

To make things worse, suddenly I started receiving emails from some of their sales team in Portuguese, I guess that my last name sounds kind of Portuguese so someone say, yeah, you write to him.

And while using their system I was not really impressed. Their cost structure was weirder than AWS (and that's saying something) and to mount a volume in an instance you had to do some funky commands.

I would NEVER trust business technology to that sort of system.

Post reply on HN