Live data from Hacker News

Microsoft Azure Outage

azure.microsoft.com

81–90 of 178 posts

Re: Microsoft Azure Outage

#81
post #77

Earlier quoted context omitted.

Does DNS propagate quickly enough to alleviate an outage or is it just a matter of ensuring that you recover within a few hours rather on waiting on an outage resolution that might take longer? Alternately, can't you just have multiple A records to distribute your load across cloud platforms and just drop the one for whichever platform is having an outage?

From experience with multi-datacenter setups, if you set a 60second TTL on your DNS records, you'll see 95%+ of traffic get the update within 5 minutes. Also, you can associate multiple addresses with a record. It's up to the client to retry on failure, but all browsers do (as far as I know)

Wouldn't that kill DNS if everyone did that considering it relies on caching for performance across the world?

Re: Microsoft Azure Outage

#82

Earlier quoted context omitted.

Over 80% of the Fortune 500 run on Azure. 20% of Azure VMs are Linux. You are not well informed.

"run on" - I suspect you're being fed a unicode pile of poo here. More likely the have _something_ which runs on Azure. Fortune 500s are, pretty much by definition, quite large - and probably have tons of departments and sub departments. And at least one of those departments probably has a task of trying out new things, like Azure, by running something on it. What surprises me is that nearly 20% of Fortune 500s _don'…

Ignoring the fortune companies, most of Microsoft's own services like Office 365 run on Azure. That's a pretty big bet right there.

Re: Microsoft Azure Outage

#83
Azure support is probably the worst one I had ever deal with. When my account (and service itself) stopped working, I haven't received any email. When I tried to sign in, all I got was some generic error saying "There's something wrong with your account". My services of course were down and I couldn't do ANYTHING. I've contacted the support to learn that my account has been blocked (!) because there was some suspicious (!!) activity going on. What the... No, they couldn't tell me what exactly it was. I've exchanged emails back and forth with the support for several days to learn nothing new, my account and services were still disabled and I was more than pissed off. From that day I hate Azure and I advise anyone against using it, because such situation is absolutely unacceptable.

Re: Microsoft Azure Outage

#84
post #32

Earlier quoted context omitted.

For what it's worth, I've been watching several Azure-hosted sites that I control and they've been coming back online sequentially (and are all back online now). Whatever they're fixing, it seems to be taking some time, but is progressing steadily at a good pace in the last hour.

Mine have been gradually coming back, too. Timing couldn't have been any better for me. Some Alanis material there: https://twitter.com/bizspark/status/534858596748906496 Now I'll have to distribute between AWS and Azure, too.

Check out Softlayer. It has been more reliable than AWS in my experience.

Re: Microsoft Azure Outage

#85
post #81
post #77

Earlier quoted context omitted.

From experience with multi-datacenter setups, if you set a 60second TTL on your DNS records, you'll see 95%+ of traffic get the update within 5 minutes. Also, you can associate multiple addresses with a record. It's up to the client to retry on failure, but all browsers do (as far as I know)

Wouldn't that kill DNS if everyone did that considering it relies on caching for performance across the world?

I'm no expert, but no. Most big sites rely on a fairly short TTL.

It's a thick layer of caches. Your browser, OS, router, ISP, and a bunch of intermediaries can cache the DNS. So even at 60s, you get good cache hits (the busier, the more true that is, of course)

Also, the update can always happen asynchronously. You and 9999 people ask your ISP for Facebook's IP. It serves all of you a slightly stale IP and asynchronously fetches a new one (thus turning 10000 requests into 1). AKA: thundering heard problem.

DNS mostly uses UDP, which is more efficient for the server and harder to DOS (the server doesn't have to maintain state per request).

Finally, # of requests is usually (always?) a factor in the price of DNS services. So the cost is borne by the clients, not the service providers. And since DNS hosting is seemingly profitable, I assume they're more than happy to build up the infrastructure to deal with additional requests.

Re: Microsoft Azure Outage

#86

Earlier quoted context omitted.

Over 80% of the Fortune 500 run on Azure. 20% of Azure VMs are Linux. You are not well informed.

"run on" - I suspect you're being fed a unicode pile of poo here. More likely the have _something_ which runs on Azure. Fortune 500s are, pretty much by definition, quite large - and probably have tons of departments and sub departments. And at least one of those departments probably has a task of trying out new things, like Azure, by running something on it. What surprises me is that nearly 20% of Fortune 500s _don'…

Agreed. When I was looking at cloud providers recently I noticed that most of them make a claim along the lines of "$X percent of {Fortune 500,FTSE 100} companies use $OURPRODUCT" where $X is > 50%. My conclusions were:-

* Most major companies use more than one cloud provider

* "Use" is a very loose term here. It could mean anything from "the accounts team in some branch office uses S3 to back up their Sage data (or uses an online backup service that uses S3 in the back end)" to "they run their main product on our infrastructure".

Re: Microsoft Azure Outage

#87
post #33

Our sites have been down for more than 3 hours now. EDIT2: Now the databases are down, this is costing us a lot of money. EDIT: Just went up again. It would be great if anyone knows how to mitigate these in the future - what can I do to protect myself against this in the future? (Except leave Azure)

Look into Cloudflare. They can act as a kind of reverse proxy to keep static stuff online. Obviously doesn't help if the transactional part of the site/database goes down, but end users will see a friendly message rather than it timing out.

Re: Microsoft Azure Outage

#88

So, the worst part about this is that zero communication has come out of Microsoft - we first started seeing issues on Sunday and filed a ticked, had an open ticket while this larger outage happened, and haven't gotten a single email saying there's an outage. I found out about it from, sigh, buzzfeed. Question - are AWS or GCE better at proactively messaging when there's an outage?

Google's operations groups are not only regularly updated during an outage, there's a root-cause analysis with remediation and prevention information posted a couple of days after any issue.

See https://groups.google.com/forum/#!forum/gce-operations and https://groups.google.com/forum/#!forum/google-appengine-dow....

Re: Microsoft Azure Outage

#89
post #33

Our sites have been down for more than 3 hours now. EDIT2: Now the databases are down, this is costing us a lot of money. EDIT: Just went up again. It would be great if anyone knows how to mitigate these in the future - what can I do to protect myself against this in the future? (Except leave Azure)

Our main cluster is on azure west us but we have another cluster on amazon east and route53 on top of that. When the main clusters fails, route53 switch to secondary, so we where not affected at all this time.

The only manual step was to delay the switch back until our vms where working fine and had all resources. We do this changing route53 health check to one that is always failing.

We had also to purge our crashed mongo nodes because the journal was broken.

https://auth0.com/availability-trust/img/auth0-infrastructur...

Re: Microsoft Azure Outage

#90

Earlier quoted context omitted.

"run on" - I suspect you're being fed a unicode pile of poo here. More likely the have _something_ which runs on Azure. Fortune 500s are, pretty much by definition, quite large - and probably have tons of departments and sub departments. And at least one of those departments probably has a task of trying out new things, like Azure, by running something on it. What surprises me is that nearly 20% of Fortune 500s _don'…

Ignoring the fortune companies, most of Microsoft's own services like Office 365 run on Azure. That's a pretty big bet right there.

Actually I don't think they do - to the best of my knowledge Office 365 didn't have any downtime as a result of this outage. And Yammer stayed up, so I assume they haven't yet migrated from AWS...
Post reply on HN