Live data from Hacker News

MS Azure down: An emerging issue is being investigated

status2.azure.com

71–80 of 131 posts

Re: MS Azure down: An emerging issue is being investigated

#71
post #61
post #48

Earlier quoted context omitted.

Do you think locally managed systems are immune to outages? Or that governments are capable of resourcing their teams sufficiently to do a better job at availability than Microsoft, Google, or Amazon?

I can think of a few good reasons. Strategic (you don’t want to hand over to a foreign power your critical infrastructures). Diversification (if all your banks run on aws, the day aws goes down you don’t have a banking system anymore). Not being at the mercy of a capricious tech billionaire (what happened to Parler could very well happen to a state if the said billionaire doesn’t like your policy).

I’m unaware of any Core Banking Systems (CBS) that runs on a cloud provider with the exception being Finastra (Azure). Other parts of retail banking stacks? Sure. Not their cores.

Re: MS Azure down: An emerging issue is being investigated

#72
post #65

Earlier quoted context omitted.

Why would Canada deploy this to a single provider? A national alert system should have a better DR plan than that.

Big cloud systems, the ones we all regularly use, get much closer scrutiny than government systems normally (because we all know when Azure, GCP, AWS, Linode, DO, CloudFlare or OVH have a major outage, misconfiguration, or A/C fire). As a "single" provider I'm sure Microsoft is much better at this than the (likely shoestring budget) government group of admins was before they moved. There's more burst capacity (if nee…

I think the point was that a national emergency system should redundantly use two cloud providers, not that they made the wrong choice of single cloud provider.

Re: MS Azure down: An emerging issue is being investigated

#73
post #48

Earlier quoted context omitted.

Do you think locally managed systems are immune to outages? Or that governments are capable of resourcing their teams sufficiently to do a better job at availability than Microsoft, Google, or Amazon?

Your comment is a fine example of the standard rhetoric from the 'move everything to the cloud' marketing people, but perhaps organizations should consider that they don't have to go with one cloud provider as a single point of failure. It's the lazy way of abstracting away responsibility and blame to some other party, and in my experience often does not result in better availability than a properly implemented "belt…

There’s a complexity cost, though, right? It’s not free to run two different systems, manage two different billing systems, different tools, pay for cross-provider bandwidth, etc.?

I’m sure there are some cases where the cost is worth the complexity, but I don’t think it’s cut and dried as written here. Most cloud providers are very reliable, and my guess is that you are more likely to have a self-inflicted outage due to the complexity of your infrastructure than the cloud provider having an outage.

Re: MS Azure down: An emerging issue is being investigated

#74

Just me or azure been really unstable lately atleast here in europe?

No, it's not just you. Independent aggregated customer data I've seen has put AWS and GCP with a marginal difference in downtime and Azure with an order of magnitude or two more downtime in comparison.

Re: MS Azure down: An emerging issue is being investigated

#75
post #65

Earlier quoted context omitted.

Why would Canada deploy this to a single provider? A national alert system should have a better DR plan than that.

Big cloud systems, the ones we all regularly use, get much closer scrutiny than government systems normally (because we all know when Azure, GCP, AWS, Linode, DO, CloudFlare or OVH have a major outage, misconfiguration, or A/C fire). As a "single" provider I'm sure Microsoft is much better at this than the (likely shoestring budget) government group of admins was before they moved. There's more burst capacity (if nee…

It's an emergency alerting system. Low bandwidth, broadcast. Perfect for blasting a few kilowatts of RF via digipeters each with their own generators and lead acid battery banks.

Re: MS Azure down: An emerging issue is being investigated

#76
post #61

Earlier quoted context omitted.

I can think of a few good reasons. Strategic (you don’t want to hand over to a foreign power your critical infrastructures). Diversification (if all your banks run on aws, the day aws goes down you don’t have a banking system anymore). Not being at the mercy of a capricious tech billionaire (what happened to Parler could very well happen to a state if the said billionaire doesn’t like your policy).

I’m unaware of any Core Banking Systems (CBS) that runs on a cloud provider with the exception being Finastra (Azure). Other parts of retail banking stacks? Sure. Not their cores.

The major cloud providers all aim for 99.999% uptime. Keep in mind that 99.99% uptime means ~4 minutes of downtime a month. I think there are other reasons that banks may not want or have the ability to run their core services on the cloud.

Re: MS Azure down: An emerging issue is being investigated

#77
post #57

Earlier quoted context omitted.

> can acheive the same high availability and elasticity Maybe not, but even just a "flick the switch to go back to a dumb system" option is worth maintaining

Is it really? You would maintain an entire parallel stack built on an entirely separate infrastructure stack, with completely different deployment patterns and all the data synced? DR is really hard if you just want to fail over to another AWS/Azure/GCP region. I can't imagine what a nightmare it would be to maintain an on-prem DR standby. To mitigate a couple hours per year (in a really bad year) of cloud downtime?

Well it depends on the system, doesn't it? For filing taxes, it doesn't matter that much if there's downtime as long as you have a plan to move the demand around (i.e. don't fine people when it's down), but for 911 calls you're going to have a lot of fairly niche infrastructure that you need to plumb in any way, so the actual cloud parts should be comparatively easy to replace.

Re: MS Azure down: An emerging issue is being investigated

#78

Just me or azure been really unstable lately atleast here in europe?

Major worldwide outages for AAD, CosmosDB, and now DNS in just the past few months so it’s definitely not just you. Maybe Microsoft should invest less in their shiny new AI/ML platforms and more into stability of their core services.

I'd like to second this. From where I'm sitting it feels like Microsoft is prioritizing marketing more than engineering.

It seems like most larger customers that use Azure do so because management got shiny presentations from Microsoft and now it's their "strategical partner".

A lot of overselling with huge discounts gets them their in. Already seen this at multiple companies. Azure is a nice platform for your Windows administrators to shift some load to the cloud. But to build large applications on?

Edit: And I kinda feel bad for saying this, since I assume that there are indeed pretty competent engineers working on Azure. But somewhere something isn't right.

Re: MS Azure down: An emerging issue is being investigated

#79

bing.com, status.azure.com, status2.azure.com - all down. Can't sign into portal.azure.com, can't hit Azure File Shares, etc. The last outage a few days ago was enough for my company to up and move most of our stuff to AWS. This new outage is enough for us to fully migrate away from Azure. What a cluster.

Instead of going from relying on a single provider to relying on a single provider, you could use both AWS and Azure.

Sure, double my budget and I’ll get right on that.

Re: MS Azure down: An emerging issue is being investigated

#80
post #34

Looks like this took down the national emergency alert system in Canada. I'm registered as an alert LMD (last mile distributor), and Pelmorex (corporation running the system) just emailed me to say "Please note that currently there is an unexpected significant outage on Microsoft Azure that is affecting the availability of the NAADS system and other clients globally. The NAAD System feeds are currently not accessible…

And here it is, the main problem with outsourcing critical infrastructure. Remote server providers like Microsoft Azure should at most be used as twins/redundant systems to a locally managed system. Governments of the world: pay your IT people more money to prevent brain drain.

[deleted]
Post reply on HN