Live data from Hacker News

Tell HN: Azure outage

news.ycombinator.com

811–820 of 841 posts

Re: Tell HN: Azure outage

#811

Earlier quoted context omitted.

It's absolutely doable if you design for it. The moment you choose to use S3 instead of hosting your own object store, though, you either use AWS because S3 and IAM already have you or spend more time on the care and feeding of your storage system as opposed to actually doing the thing you customers are paying you to do. It's not impossible, just complicated and difficult for any moderately complex architecture.

There are plenty of compatible S3-like offerings. That's one of the lesser things that tie me to a cloud.

Even on non-AWS projects, I still use S3. I haven't really explored the other options, but if you have opinions or advice I'd love to hear them.

One thing very important, is that I can authorise specific web clients (users) to access specific resources from S3. Such as a document that he can download, but others with the link should not be able to download.

Thank you!

Re: Tell HN: Azure outage

#815

Earlier quoted context omitted.

Not sure how the current situation is better. Being stranded with no way whatsoever to access most/all of your services sounds way more terrifying than regular issues limited to a couple of services at a time

> no way whatsoever to access most/all of your services I work on a product hosted on Azure. That's not the case. Except for front door, everything else is running fine. (Front door is a reverse proxy for static web sites.) The product itself (an iot stormwater management system) is running, but our customers just can't access the website. If they need to do something, they can go out to the sites or call us and we c…

The portal was down for most of the day and accessing any resources from the portal once the portal was up was not possible.

Re: Tell HN: Azure outage

#816
post #768
post #708

Earlier quoted context omitted.

In Microsoft's defense, Azure has always been a complete joke. It's extremely developer unfriendly, buggy and overpriced.

> In Microsoft's defense, Azure has always been a complete joke. It's extremely developer unfriendly, buggy and overpriced. Don't forget extremely insecure. There is a quarterly critical cross-tenant CVE with trivial exploitation for them, and it has been like that for years.

Given how much time I spent on my first real multi-tenant project, dealing with the consequences of architecture decisions meant to prevent these sorts of issues, I can see clearly the temptation to avoid dealing with them.

But what we do when things are easy is not who we are. That's a fiction. It's how we show up when we are in the shit that matters. It's discipline that tells you to voluntarily go into all of the multi-tenant mitigations instead of waiting for your boss to notice and move the goalposts you should have moved on your own.

Re: Tell HN: Azure outage

#817

Earlier quoted context omitted.

33 minutes from impact to status page for a complete outage is a joke.

As a technologist, you should always avoid MS. Even if they have a best-in-class solution for some domain, they will use that to leverage you into their absolute worst-in-class ecosystem.

I see Amazon using a subset of the same sorts of obfuscations that Microsoft was infamous for. They just chopped off the crusts so it's less obvious that it's the same shit sandwich.

Re: Tell HN: Azure outage

#818

Earlier quoted context omitted.

33 minutes from impact to status page for a complete outage is a joke.

More importantly `15:45 UTC on 29 October 2025 – Customer impact began. 16:04 UTC on 29 October 2025 – Investigation commenced following monitoring alerts being triggered. ` A 19-minute delay in alert is a joke.

10 minutes to alert, to avoid flapping false positives. 10 minute response window for first responders. Or, 5 minute window before failing over to backup alerts, and 4 minutes to wake up, have coffee, and open the appropriate windows.

Re: Tell HN: Azure outage

#819
post #770

Earlier quoted context omitted.

More importantly `15:45 UTC on 29 October 2025 – Customer impact began. 16:04 UTC on 29 October 2025 – Investigation commenced following monitoring alerts being triggered. ` A 19-minute delay in alert is a joke.

That does not say it took 19 minutes for alerts to appear. Following could mean any amount of time.

It's 19 minutes until active engagement by staff. And planned rolling restarts can trigger alerts if you don't set thresholds of time instead of just thresholds of count.

It would be nice though if alert systems made it easy to wire up CD to turn down sensitivity during observed actions. Sort of like how the immune system turns down a bit while you're eating.

Re: Tell HN: Azure outage

#820
post #677

Earlier quoted context omitted.

Is it possible to trace your own vote after? There has to be a technical solution to ensure that your own vote was counted

yes there is. Check double envelope mail in voting mechanics.

That's just that they got my ballot. How to ensure they allocated my specific vote to the specific candidate/measure.
Post reply on HN