Live data from Hacker News

Microsoft Azure Outage

twitter.com

201–210 of 247 posts

Re: Microsoft Azure Outage

#201

Earlier quoted context omitted.

But shouldn't the individual service dots be automatically turning another color than green? I mean it's an automated service status page, right? Whether there is a human message at the top and that can take some time I understand.

No, it's not automated. I'm sure the underlying tech is automated, but once companies grow beyond a certain size, it needs a human to say "show this status change to the world" because there are lots of things depending on it (e.g. SLAs, but also bonuses, I assume), so they don't want a potential bug in the status system to influence that. It's weird how slow they are with manual sign-off though.

But I don’t want the page connected to their bonuses or SLA’s I just want to know whether they are having any issues anywhere. And I need to know within a minute of my own service not working so I’m not chasing the wrong thing. This can’t be an unreasonable thing to ask for?

Re: Microsoft Azure Outage

#202
post #97
post #84

Shows that all these availability zones and regions don't really help if an outage can knock out a whole cloud provider. And that's not specific to Microsoft. The only way to really ensure uptime is to use two providers. Sadly, that's basically only possible with on-prem/colocation where traffic is cheap.

It's mostly Azure though that is badly designed to such an extent that multiple times there have been global outages. In general Azure availability, security (the only major cloud provider with not one but multiple cross-tenant security exploits) and usability are pretty terrible so it shouldn't be used for anything but saying "this is how it should not be done". GCP had a similar thing once, where a BGP update knock…

To be fair, AWS once had a global Route53 outage, which was effectively a global outage for anyone using AWS for DNS.

Re: Microsoft Azure Outage

#205

What's the point of having a status page if it doesn't indicate the issues? https://status.azure.com/en-us/status Azure, Teams, Outlook are almost down from Greece and Germany, and their status page shows that everything is fine :-)

[deleted]

Re: Microsoft Azure Outage

#206
post #202
post #97

Earlier quoted context omitted.

It's mostly Azure though that is badly designed to such an extent that multiple times there have been global outages. In general Azure availability, security (the only major cloud provider with not one but multiple cross-tenant security exploits) and usability are pretty terrible so it shouldn't be used for anything but saying "this is how it should not be done". GCP had a similar thing once, where a BGP update knock…

To be fair, AWS once had a global Route53 outage, which was effectively a global outage for anyone using AWS for DNS.

Do you have a link to an article about that? My google-fu is weak, and this sounds interesting - that should not happen to DNS - at all - and from the outside Route53 looks quite well managed. So what the heck did they do?

Re: Microsoft Azure Outage

#207
post #174

Earlier quoted context omitted.

It has nothing to do with press. This is negative press already, and journalist can use this to write their stories without waiting for the official light to go from green to yellow. It's about contractual obligations and SLAs. Things are not officially down in most agreements until MSFT acknowledges they're down. Refunds issued because your blob storage failed to meet 99.9999 uptime to your largest customers are dir…

I'm not going out of my way to be hyperbolic or anything here, but that sounds suspiciously like "fraud" to me.

I don't think they're committing fraud.

I think it's an important enough page that it can't be automated. It needs a manual approval from a human, for the very basics, like even if the status reporting system is operating correctly, because of various downstream effects.

Re: Microsoft Azure Outage

#208

Earlier quoted context omitted.

No, it's not automated. I'm sure the underlying tech is automated, but once companies grow beyond a certain size, it needs a human to say "show this status change to the world" because there are lots of things depending on it (e.g. SLAs, but also bonuses, I assume), so they don't want a potential bug in the status system to influence that. It's weird how slow they are with manual sign-off though.

But I don’t want the page connected to their bonuses or SLA’s I just want to know whether they are having any issues anywhere. And I need to know within a minute of my own service not working so I’m not chasing the wrong thing. This can’t be an unreasonable thing to ask for?

I agree. I'm already annoyed at Hetzner with their 5 minute lag in reporting network outages where I'm regularly noticing them, investigating, checking status and then only after a few minutes see them updating and saying "it's us".

If you work with Microsoft, you might as well spend a few bucks extra and have an external monitoring system monitor Microsoft's systems so you get real-time third-party confirmation when your monitoring alerts you of issues concerning your system. It's the price you pay for scale, I guess. More money involved = more lawyers involved = more accountants involved = more MBAs involved = more corporate bullshit.

Re: Microsoft Azure Outage

#210

Earlier quoted context omitted.

Leave to go where? On-premise and being miserable having to wait months to get a new server with poor automation, observability and worse outages? To another major cloud provider with similar pricing and outages? Cloud helped mostly with automation and scaling but if your system is that critical, you should consider a good CDN as load balancer and multi-cloud (or at least multi-region) for actual robustness.

AWS and GCP both have ~100% uptime in every region for VMs this month. Meanwhile the majority of Azure regions have had various outages in the same period: https://cloudharmony.com/status-of-compute

Wow I didn't expect the difference to be so obvious.
Post reply on HN