Live data from Hacker News

Tell HN: Azure outage

news.ycombinator.com

351–360 of 841 posts

Re: Tell HN: Azure outage

#352

Earlier quoted context omitted.

It's red right now.

Only for the Azure Portal, despite Front Door also being down but showing as green on the status page.

Heh, now it says Front Door and "Network Infrastructure" are down. That second one seems bad.

Re: Tell HN: Azure outage

#353
post #33

Update 16:57 UTC: Azure Portal Access Issues Starting at approximately 16:00 UTC, we began experiencing Azure Front Door issues resulting in a loss of availability of some services. In addition. customers may experience issues accessing the Azure Portal. Customers can attempt to use programmatic methods (PowerShell, CLI, etc.) to access/utilize resources if they are unable to access the portal directly. We have faile…

They briefly had a statement about using Traffic Manager to work with your AFD to work around this issue, with a link to learn.microsoft.com/...traffic-manager, and the link didn't work. Due to the same issue affecting everyone right now.

They quickly updated the message to REMOVE the link. Comical at this point.

Re: Tell HN: Azure outage

#354
post #166

Earlier quoted context omitted.

It's pretty unlikely. AWS published a public 'RCA' https://aws.amazon.com/message/101925/ . A race condition in a DNS 'record allocator' causing all DNS records for DDB to be wiped out. I'm simplifying a bit, but I don't think it's likely that Azure has a similar race condition wiping out DNS records on _one_ system than then propagates to all others. The similarity might just end at "it was DNS".

It's always DNS

It is a coin flip, heads DNS, tails BGP

Re: Tell HN: Azure outage

#355
I was working when I saw the portal page showing only resource groups and lots of items missing. I thought it was a weird browser cache issue.

The actual stuff I was working on (App Insights, Function App) that was still open was operational.

Re: Tell HN: Azure outage

#358

The sad thing is - $MSFT isn't even down by 1%. And IIRC, $AMZN actually went up during their previous outage. So if we look at these companies' bottom lines, all those big wigs are actually doing something right. Sales and lobbying capacity is way more effective than reliability or good engineering (at least in the short term).

Look how important we are, is what these failures show

Re: Tell HN: Azure outage

#360

Earlier quoted context omitted.

Ahh, but you forget what it used to be like. Sites used to go down all the time. Now, they go down a lot less frequently, but when they do, it's more widespread.

Not sure how the current situation is better. Being stranded with no way whatsoever to access most/all of your services sounds way more terrifying than regular issues limited to a couple of services at a time

> no way whatsoever to access most/all of your services

I work on a product hosted on Azure. That's not the case. Except for front door, everything else is running fine. (Front door is a reverse proxy for static web sites.)

The product itself (an iot stormwater management system) is running, but our customers just can't access the website. If they need to do something, they can go out to the sites or call us and we can "rub two sticks together" and bypass the website. (We could also bypass front door if someone twisted our arms.)

Most customers only look at the website a few times a year.

---

That being said, our biggest point of failure is a completely different iot vendor who you probably won't hear about on Hacker News when they, or their data networks, have downtime.

Post reply on HN