Preliminary post incident review: https://azure.status.microsoft/en-gb/status/history/ Timeline 15:45 UTC on 29 October 2025 – Customer impact began. 16:04 UTC on 29 October 2025 – Investigation commenced following monitoring alerts being triggered. 16:15 UTC on 29 October 2025 – We began the investigation and started to examine configuration changes within AFD. 16:18 UTC on 29 October 2025 – Initial communication po…
Tell HN: Azure outage
701–710 of 841 posts
Re: Tell HN: Azure outage
#702Earlier quoted context omitted.
The "Blades" experience [0] where instead of navigating between pages it just kept opening things to the side and expanding horizontally? Yeah, that had some fun ideas but was way more confusing than it needed to be. But also that was quite a few years back now. The Portal ditched that experience relatively quickly. Just long enough to leave a lot of awful first impressions, but not long enough for it to be much more…
Azure to me has always suffered from a belief that “UI innovations can solve UX complexity if you just try hard enough.” Like, AWS, and GCP to a lesser extent, has a principled approach where simple click-ops goals are simple. You can access the richer metadata/IAM object model at any time, but the wizards you see are dumb enough to make easy things easy. With Azure, those blades allow tremendously complex “you need…
I don't want to pay for or lock myself into, "Azure Insights".
I just want to see the logging, that I know if I can remember the right buttons to click, are available.
The worst place to try is "Monitoring > Logs", this is where you get faced up front with a query designer. I've never worked out how to do a simple "list by time" on that query designer, but it doesn't matter, because if you suffer through that UX, you find out that's not actually where the logs are anyway.
You have to go down a different path. Don't be distracted by "Log Stream", that's not it either, it sounds useful but it's not. By default it doesn't log anything. If you do configure it to log, then it still doesn't actually log everything.
What you have to actually do, and I've had to open the portal to check this, is click "Diagnose and Solve Problems" and then look for "Diagnostic tools" and then a small link to "Application Event Logs".
Finally you get to your logs, although it's still a bad way to try to view logs, it's at least marginally better than the real windows event viewer, an application that feels like it hasn't been updated since NT4. ( Although some might suggest that's a good thing. )
Re: Tell HN: Azure outage
#703Earlier quoted context omitted.
We're multi-cloud and it really saved a few workloads last week with the AWS issue. It's not easy though.
This is the eternal tension for early-stage builders, isn't it? Multi-cloud gives you resilience, but adds so much complexity that it can actually slow down shipping features and iterating. I'm curious—at what point did you decide the overhead was worth it? Was it after experiencing an outage, or did you architect for it from day one? As someone launching a product soon (more on the builder/product side than infra-en…
I would recommend focusing on multi-region within a single CSP instead (both for workloads AND your tooling), which covers the vast majority of incidents and lays some of the architectural foundation for multi-cloud down the road. Develop failover plans for each service in your architecture (eg. planned/tested runbooks to migrate to Traffic Manager in the event AFD goes down)
Also choose your provider wisely. We experience 3-5x the number of service-impacting incidents on Azure that we do on AWS. I'm sure others have different experiences, but I would never personally start a company on Azure. AWS has its own issues, of course, but reliability has not been a major one (relatively speaking) over the past 10 years. Last week's incident with DynamoDB in us-east-1 had zero impact on our AWS workloads in other regions.
Re: Tell HN: Azure outage
#704It still surprises me how much essential services like public transport are completely reliant on cloud providers, and don't seem to have backups in place. Here in The Netherlands, almost all trains were first delayed significantly, and then cancelled for a few hours because of this, which had real impact because today is also the day we got to vote for the next parlement (I know some who can't get home in time befor…
Re: Tell HN: Azure outage
#705Currently standing in a half closed supermarket because the tills are down and they cant take payments
Re: Tell HN: Azure outage
#706For some reason an Azure outage does not faze me in the same way that an AWS outage does. I have never had much confidence in Azure as a cloud provider. The vertical integration of all the things for a Microsoft shop was initially very compelling. I was ready to fight that battle. But, this fantasy was quickly ruined by poor execution on Microsoft's part. They were able to convince me to move back to AWS by simply ma…
The only reason you'd notice MS was down was if Github was down....
Re: Tell HN: Azure outage
#707Preliminary post incident review: https://azure.status.microsoft/en-gb/status/history/ Timeline 15:45 UTC on 29 October 2025 – Customer impact began. 16:04 UTC on 29 October 2025 – Investigation commenced following monitoring alerts being triggered. 16:15 UTC on 29 October 2025 – We began the investigation and started to examine configuration changes within AFD. 16:18 UTC on 29 October 2025 – Initial communication po…
At 16:04 “Investigation commenced”. Then at 16:15 “We began the investigation”. Which is it?
Re: Tell HN: Azure outage
#708Preliminary post incident review: https://azure.status.microsoft/en-gb/status/history/ Timeline 15:45 UTC on 29 October 2025 – Customer impact began. 16:04 UTC on 29 October 2025 – Investigation commenced following monitoring alerts being triggered. 16:15 UTC on 29 October 2025 – We began the investigation and started to examine configuration changes within AFD. 16:18 UTC on 29 October 2025 – Initial communication po…
33 minutes from impact to status page for a complete outage is a joke.
Re: Tell HN: Azure outage
#709Earlier quoted context omitted.
Why? Starbucks is not providing a critical service. Spending less money and resources and just accepting the risk that occasionally you won't be able to sell coffee for a few hours is a completely valid decision from both management and engineering pov.
Or maybe we should throw them in jail.
Re: Tell HN: Azure outage
#710Earlier quoted context omitted.
At 16:04 “Investigation commenced”. Then at 16:15 “We began the investigation”. Which is it?
Quick coffee run before we get stuck in mate
You don’t want to debug stuff with low sugar.