Live data from Hacker News

MS Azure down: An emerging issue is being investigated

status2.azure.com

91–100 of 131 posts

Re: MS Azure down: An emerging issue is being investigated

#91

Earlier quoted context omitted.

I’m unaware of any Core Banking Systems (CBS) that runs on a cloud provider with the exception being Finastra (Azure). Other parts of retail banking stacks? Sure. Not their cores.

The major cloud providers all aim for 99.999% uptime. Keep in mind that 99.99% uptime means ~4 minutes of downtime a month. I think there are other reasons that banks may not want or have the ability to run their core services on the cloud.

[deleted]

Re: MS Azure down: An emerging issue is being investigated

#92

Edit: it seems like it might be over. So, the outage appears to be that DNS for `azure.com` and maybe also `windows.net` (blob storage for us, but I'm not sure) is not resolving. So, the OP's link here is broken. Tweets indicate that it might be intermittently resolving. https://twitter.com/AzureSupport/status/1377737333307437059 > Warning sign We are aware of an issue affecting the Azure Portal and Azure services, p…

It was less than a month ago and, with us starting our migration to Azure which is due to complete by mid-June, I'm getting decidedly jumpy about the amount of downtime.

Our current provider, with whom we have a number of dedicated servers, might be piss poor in some ways (takes ages to get changes made, lack of pricing transparency, kind of expensive for what they are), but I don't remember the last time they had an outage. There's definitely been one in the last year, but two in the space of a month is ridiculous.

Is this normal for Azure? Is AWS any better?

Re: MS Azure down: An emerging issue is being investigated

#93

Earlier quoted context omitted.

AWS: https://downdetector.com/status/aws-amazon-web-services/ Cloudflare: https://downdetector.com/status/cloudflare/

The problem with downdetector is that people say something is down when really it's another service. Like with Cloudflare, a lot of the comments are simply that a website was down giving an CF error, but in reality it was probably not CF that was down but an underlying service.

On the other hand, CF is themselves known for blaming the underlying site in their error pages when in fact the underlying wasn't the problem.

So it's just a mess.

Re: MS Azure down: An emerging issue is being investigated

#94
post #65

Earlier quoted context omitted.

Big cloud systems, the ones we all regularly use, get much closer scrutiny than government systems normally (because we all know when Azure, GCP, AWS, Linode, DO, CloudFlare or OVH have a major outage, misconfiguration, or A/C fire). As a "single" provider I'm sure Microsoft is much better at this than the (likely shoestring budget) government group of admins was before they moved. There's more burst capacity (if nee…

I think the point was that a national emergency system should redundantly use two cloud providers, not that they made the wrong choice of single cloud provider.

The national emergency system doesn't rely on one provider[0] or one cloud, while Pelmorex Corp (The Weather Network) is part of the chain (and a curious one) it isn't the entire system, nor is their choice of provider/hardware a government choice (but it doesn't seem poor). The network isn't down, an aspect may have endured an outage. For significant (but not total) coverage of the country the "Alert Distributors" could be covered with contacting 15 corporations (huzzah anti competitive Canada)... or less if you pick a specific channel (eg Wireless=4), which I imagine is part of alerting protocol (your cable, mobile, radio, TV, and web applications don't all send the same alert)

[0]: https://www.publicsafety.gc.ca/cnt/mrgnc-mngmnt/mrgnc-prprdn...

Re: MS Azure down: An emerging issue is being investigated

#97

Earlier quoted context omitted.

Your comment is a fine example of the standard rhetoric from the 'move everything to the cloud' marketing people, but perhaps organizations should consider that they don't have to go with one cloud provider as a single point of failure. It's the lazy way of abstracting away responsibility and blame to some other party, and in my experience often does not result in better availability than a properly implemented "belt…

Or it’s borne of experience of doing it oneself being much less reliable than cloud providers even accounting for these rare outages.

[deleted]

Re: MS Azure down: An emerging issue is being investigated

#98

Earlier quoted context omitted.

AWS: https://downdetector.com/status/aws-amazon-web-services/ Cloudflare: https://downdetector.com/status/cloudflare/

The problem with downdetector is that people say something is down when really it's another service. Like with Cloudflare, a lot of the comments are simply that a website was down giving an CF error, but in reality it was probably not CF that was down but an underlying service.

I assume downdetector can isolate region, run a traceroute to at least at some extent analyse source of network issues?

Re: MS Azure down: An emerging issue is being investigated

#100

Edit: it seems like it might be over. So, the outage appears to be that DNS for `azure.com` and maybe also `windows.net` (blob storage for us, but I'm not sure) is not resolving. So, the OP's link here is broken. Tweets indicate that it might be intermittently resolving. https://twitter.com/AzureSupport/status/1377737333307437059 > Warning sign We are aware of an issue affecting the Azure Portal and Azure services, p…

It was less than a month ago and, with us starting our migration to Azure which is due to complete by mid-June, I'm getting decidedly jumpy about the amount of downtime. Our current provider, with whom we have a number of dedicated servers, might be piss poor in some ways (takes ages to get changes made, lack of pricing transparency, kind of expensive for what they are), but I don't remember the last time they had an…

We’ve been with Azure for two years. This scale of an issue is definitely abnormal, but other minor outages (typically geo-specific) are more frequent, which is why we do have some things using their “paired regions” geo-redundancy mechanisms.
Post reply on HN