Live data from Hacker News

MS Azure down: An emerging issue is being investigated

status2.azure.com

61–70 of 131 posts

Re: MS Azure down: An emerging issue is being investigated

#61
post #48

Earlier quoted context omitted.

And here it is, the main problem with outsourcing critical infrastructure. Remote server providers like Microsoft Azure should at most be used as twins/redundant systems to a locally managed system. Governments of the world: pay your IT people more money to prevent brain drain.

Do you think locally managed systems are immune to outages? Or that governments are capable of resourcing their teams sufficiently to do a better job at availability than Microsoft, Google, or Amazon?

I can think of a few good reasons. Strategic (you don’t want to hand over to a foreign power your critical infrastructures). Diversification (if all your banks run on aws, the day aws goes down you don’t have a banking system anymore). Not being at the mercy of a capricious tech billionaire (what happened to Parler could very well happen to a state if the said billionaire doesn’t like your policy).

Re: MS Azure down: An emerging issue is being investigated

#63
post #48

Earlier quoted context omitted.

And here it is, the main problem with outsourcing critical infrastructure. Remote server providers like Microsoft Azure should at most be used as twins/redundant systems to a locally managed system. Governments of the world: pay your IT people more money to prevent brain drain.

Do you think locally managed systems are immune to outages? Or that governments are capable of resourcing their teams sufficiently to do a better job at availability than Microsoft, Google, or Amazon?

The comment you're replying to says the cloud should be used as a redundant system. So they are inherently saying they are not immune to outages.

Re: MS Azure down: An emerging issue is being investigated

#65
post #34

Looks like this took down the national emergency alert system in Canada. I'm registered as an alert LMD (last mile distributor), and Pelmorex (corporation running the system) just emailed me to say "Please note that currently there is an unexpected significant outage on Microsoft Azure that is affecting the availability of the NAADS system and other clients globally. The NAAD System feeds are currently not accessible…

Why would Canada deploy this to a single provider? A national alert system should have a better DR plan than that.

Big cloud systems, the ones we all regularly use, get much closer scrutiny than government systems normally (because we all know when Azure, GCP, AWS, Linode, DO, CloudFlare or OVH have a major outage, misconfiguration, or A/C fire). As a "single" provider I'm sure Microsoft is much better at this than the (likely shoestring budget) government group of admins was before they moved. There's more burst capacity (if needed) without paying for redundant equipment most of the year, and ultimately it's probably at a lower cost.

Meanwhile Azure is the only cloud with two Canada regions (Canada "Central", Canada East)[0], AWS has one (Canada "Central")[1], GCP has one (Montreal same city as AWS)[2].. there's really only one player if you want to use managed cloud services.

[0]: https://azure.microsoft.com/en-us/global-infrastructure/geog... [1]: https://aws.amazon.com/about-aws/global-infrastructure/regio... [2]: https://cloud.google.com/about/locations/

Re: MS Azure down: An emerging issue is being investigated

#66
post #34

Looks like this took down the national emergency alert system in Canada. I'm registered as an alert LMD (last mile distributor), and Pelmorex (corporation running the system) just emailed me to say "Please note that currently there is an unexpected significant outage on Microsoft Azure that is affecting the availability of the NAADS system and other clients globally. The NAAD System feeds are currently not accessible…

And here it is, the main problem with outsourcing critical infrastructure. Remote server providers like Microsoft Azure should at most be used as twins/redundant systems to a locally managed system. Governments of the world: pay your IT people more money to prevent brain drain.

I'd say that you need critical infrastructure redundantly deployed everywhere, and experiencing a training outage on one of the platform each quarter.

But wanting is one thing, having the money to implement and maintain such a solution, a quite another thing :-/

Re: MS Azure down: An emerging issue is being investigated

#67
post #55

> Microsoft rerouted traffic to our resilient DNS capabilities and are seeing improvement in service availability At least use well formed English when posting your update...

You're not on call very often, are you?

I suspect that the update started as “We” (which was grammatically correct) and got hastily changed along the way to Microsoft.

Re: MS Azure down: An emerging issue is being investigated

#68
post #55

> Microsoft rerouted traffic to our resilient DNS capabilities and are seeing improvement in service availability At least use well formed English when posting your update...

You're not on call very often, are you?

Would I forget English if I was?

To be clear: I’m not ridiculing anyone in particular (native or non-native). I’m pointing out that a multi-billion dollar company experiencing a major outage could probably proof read the statements they put out about that incident before publishing. Else why not just go “if ur dns is bad its probs us soz, were on it”

Re: MS Azure down: An emerging issue is being investigated

#69

Earlier quoted context omitted.

AWS: https://downdetector.com/status/aws-amazon-web-services/ Cloudflare: https://downdetector.com/status/cloudflare/

DownDetector is worthless. There’s no reason to put any stock in it. Look at the comments on any entry and it’s clear people do stuff like report “outages” for Google because a random website won’t work in Chrome.

You could also argue that it is still better than always-green status pages of cloud providers.

Re: MS Azure down: An emerging issue is being investigated

#70
post #57
post #51

Earlier quoted context omitted.

So you think local IT can acheive the same high availability and elasticity? Sorry that isn't usually my experience. Lots of anecdotal local IT get lucky, but on average I think this is the wrong lesson to learn.

> can acheive the same high availability and elasticity Maybe not, but even just a "flick the switch to go back to a dumb system" option is worth maintaining

Is it really? You would maintain an entire parallel stack built on an entirely separate infrastructure stack, with completely different deployment patterns and all the data synced? DR is really hard if you just want to fail over to another AWS/Azure/GCP region. I can't imagine what a nightmare it would be to maintain an on-prem DR standby. To mitigate a couple hours per year (in a really bad year) of cloud downtime?
Post reply on HN