For some reason an Azure outage does not faze me in the same way that an AWS outage does. I have never had much confidence in Azure as a cloud provider. The vertical integration of all the things for a Microsoft shop was initially very compelling. I was ready to fight that battle. But, this fantasy was quickly ruined by poor execution on Microsoft's part. They were able to convince me to move back to AWS by simply ma…
Tell HN: Azure outage
681–690 of 841 posts
Re: Tell HN: Azure outage
#682Timeline
15:45 UTC on 29 October 2025 – Customer impact began.
16:04 UTC on 29 October 2025 – Investigation commenced following monitoring alerts being triggered.
16:15 UTC on 29 October 2025 – We began the investigation and started to examine configuration changes within AFD.
16:18 UTC on 29 October 2025 – Initial communication posted to our public status page.
16:20 UTC on 29 October 2025 – Targeted communications to impacted customers sent to Azure Service Health.
17:26 UTC on 29 October 2025 – Azure portal failed away from Azure Front Door.
17:30 UTC on 29 October 2025 – We blocked all new customer configuration changes to prevent further impact.
17:40 UTC on 29 October 2025 – We initiated the deployment of our ‘last known good’ configuration.
18:30 UTC on 29 October 2025 – We started to push the fixed configuration globally.
18:45 UTC on 29 October 2025 – Manual recovery of nodes commenced while gradual routing of traffic to healthy nodes began after the fixed configuration was pushed globally.
23:15 UTC on 29 October 2025 - PowerApps mitigation of dependency, and customers confirm mitigation.
00:05 UTC on 30 October 2025 – AFD impact confirmed mitigated for customers.
Re: Tell HN: Azure outage
#683Wow, they are still down 12 hours later. :/
Re: Tell HN: Azure outage
#684Re: Tell HN: Azure outage
#685Earlier quoted context omitted.
i'm not sure this is an easily solvable problem. i remember reading an article arguing that your cloud provider is part of your tech stack and it's close to impossible/a huge PITA to make a non-trivial service provider-agnostic. they'd have to run their own openstack in different datacenters, which would be costly and have their own points of failure.
I run non trivial services on EC2, using that service as a VPS. My deploy script works just as well on provisioned Digital Ocean services and on docker containers using docker-compose. I do need a human to provision a few servers and configure e.g. load balancing and when to spin up additional servers under load. But that is far less of a PITA than having my systems tied to a specific provider or down whenever a clou…
Re: Tell HN: Azure outage
#686Earlier quoted context omitted.
If India can have voters vote and tally all the votes in one day, then so can everyone else. It’s the best way to avoid fraud and people going with whoever is ahead. I am sympathetic with emergency protocols for deadly pandemics, but for all else, in-person on a given day.
If it's not a national holiday where the vast majority of people don't have to work, and if there aren't polling places reasonably near every voting age citizen, it's voter suppression.
Re: Tell HN: Azure outage
#687Preliminary post incident review: https://azure.status.microsoft/en-gb/status/history/ Timeline 15:45 UTC on 29 October 2025 – Customer impact began. 16:04 UTC on 29 October 2025 – Investigation commenced following monitoring alerts being triggered. 16:15 UTC on 29 October 2025 – We began the investigation and started to examine configuration changes within AFD. 16:18 UTC on 29 October 2025 – Initial communication po…
Re: Tell HN: Azure outage
#688Re: Tell HN: Azure outage
#689Quite close to the recent AWS outage. Let me take a look if its a major one similar to AWS. Any guess on what's causing it? In hindsight, I guess the foresight of some organizations to go multi-cloud was correct after all.
We're multi-cloud and it really saved a few workloads last week with the AWS issue. It's not easy though.
I'm curious—at what point did you decide the overhead was worth it? Was it after experiencing an outage, or did you architect for it from day one?
As someone launching a product soon (more on the builder/product side than infra-engineer), I keep wrestling with this. The pragmatist in me says "start simple, prove the concept, then layer in resilience." But then you see events like this week and think "what if this happens during launch?"
How did you handle the operational complexity? Did you need dedicated DevOps folks, or are there patterns/tools that made it manageable for a smaller team?
Re: Tell HN: Azure outage
#690Earlier quoted context omitted.
CSAM apparently also means Customer Success Account Manager for those who might have gotten startled by this message like me.
Alternative für Deutschland was strange enough, when I saw CSAM I was really wondering what thread I had stumbled into