Live data from Hacker News

Is Northern Virginia still the least reliable AWS region?

statusgator.com

61–70 of 84 posts

Re: Is Northern Virginia still the least reliable AWS region?

#61
post #59

I think if you need something more reliable than us-east-1 that you should be hosting on prem in facilities you own and operate. There aren't that many businesses that truly can't handle the worst case (so far) AWS outage. Payment processing is the strongest example I can come up with that is incompatible with the SLA that a typical cloud provider can offer. Visa going down globally for even a few minutes might be wo…

> worst case (so far)

It’s kind of amazing that after nearly 20 years of “cloud”, the worst case so far still hasn’t been all that bad. Outages are the mildest type of incident. A true cloud disaster would be something like a major S3 data loss event, or a compromise of the IAM control plane. That’s what it would take for people to take multi-region/multi-cloud seriously.

Re: Is Northern Virginia still the least reliable AWS region?

#62
post #61
post #59

I think if you need something more reliable than us-east-1 that you should be hosting on prem in facilities you own and operate. There aren't that many businesses that truly can't handle the worst case (so far) AWS outage. Payment processing is the strongest example I can come up with that is incompatible with the SLA that a typical cloud provider can offer. Visa going down globally for even a few minutes might be wo…

> worst case (so far) It’s kind of amazing that after nearly 20 years of “cloud”, the worst case so far still hasn’t been all that bad. Outages are the mildest type of incident. A true cloud disaster would be something like a major S3 data loss event, or a compromise of the IAM control plane. That’s what it would take for people to take multi-region/multi-cloud seriously.

> A true cloud disaster would be something like a major S3 data loss event

So like the OVH data center fire back in 2021?

Re: Is Northern Virginia still the least reliable AWS region?

#63
post #62
post #61

Earlier quoted context omitted.

> worst case (so far) It’s kind of amazing that after nearly 20 years of “cloud”, the worst case so far still hasn’t been all that bad. Outages are the mildest type of incident. A true cloud disaster would be something like a major S3 data loss event, or a compromise of the IAM control plane. That’s what it would take for people to take multi-region/multi-cloud seriously.

> A true cloud disaster would be something like a major S3 data loss event So like the OVH data center fire back in 2021?

No, a major one.

(No shade on OVH, but they are ~1% market share player)

Re: Is Northern Virginia still the least reliable AWS region?

#64
post #61
post #59

I think if you need something more reliable than us-east-1 that you should be hosting on prem in facilities you own and operate. There aren't that many businesses that truly can't handle the worst case (so far) AWS outage. Payment processing is the strongest example I can come up with that is incompatible with the SLA that a typical cloud provider can offer. Visa going down globally for even a few minutes might be wo…

> worst case (so far) It’s kind of amazing that after nearly 20 years of “cloud”, the worst case so far still hasn’t been all that bad. Outages are the mildest type of incident. A true cloud disaster would be something like a major S3 data loss event, or a compromise of the IAM control plane. That’s what it would take for people to take multi-region/multi-cloud seriously.

I mean, EBS went offline and people were ok to continue using AWS…

https://arstechnica.com/information-technology/2011/04/amazo...

Re: Is Northern Virginia still the least reliable AWS region?

#66
post #56

Earlier quoted context omitted.

I like this a lot, this is a great comparison for hetzner american offerings since it's not big enough for them to even bother investing much into it so there's not that many complains about it. People just dumping it (me included) after discovering the amount of random issues it has probably also doesn't help. if you are using hetzner: avoid everything other than fra region, ideally pray that you are placed in the n…

Hetzner does not have any "fra region". They have Helsinki, Falkstien and Nuremberg in Europe. None of them which has any issues as far as I know. They used to have some issues with the very old stuff in Falkstien.

sorry, fsn* I have them typod fra internally and keep messing it up since it's stuck in my head.

Re: Is Northern Virginia still the least reliable AWS region?

#67
This story missed a glaring detail. There are simply more data centers in northern VA [0]. More than the rest of the US by a wide margin, or the entire EU+Asia. Things break here because it's where most things are.

[0]: https://www.datacenters.com/providers/amazon-aws/data-center...

Re: Is Northern Virginia still the least reliable AWS region?

#68
post #27

Earlier quoted context omitted.

IIRC, some AWS services are solely deployed on and/or entirely dependent on us-east-1. I don't recall which ones, but I very distinctly remember this coming up once.

IAM and Route53 have dependencies on us-east-1. AWS Organizations/Account management is us-east-1. And if you want a CDN with a custom hostname and want TLS…you have to use us-east-1.

The Route53 control plane is in us-east-1, with an optional temporary auto-failover to us-west-2 during outages. The data plane for public zones is globally distributed and highly resilient, with a 100% SLA. It continues to serve DNS records during regular control plane outages in us-east-1, but access to make changes is lost during outages.

CloudFront CDN has a similar setup. The SSL certificate and key have to be hosted in us-east-1 for control plane operations but once deployed, the public data plane is globally or regionally dispersed. There is no auto failover for the cert dependency yet. The SLA is only three 9s. Also depends on Route53.

The elephant in the room for hyperscalers is the potential for rogue employees or a cyber attack on a control plane. Considering the high stakes and economic criticality of these platforms, both are inevitable and both have likely already happened.

Re: Is Northern Virginia still the least reliable AWS region?

#69
post #60

Earlier quoted context omitted.

Bizarre way of making decisions. us-east-2 is objectively a better region to pick if you want US east, yet you feel safer picking use1 because “I’m safer making a worse decision that everyone understands is worse, as long as everyone else does it as well.”

If my cloud provider goes down and my site is offline, my customers and my boss will be upset with me and demand I fix it as fast as possible. They will not care what caused it. If my cloud provider goes down and also takes down Spotify, Snapchat, Venmo, Reddit, and a ton of other major services that my customers and my boss use daily, they will be much more understanding that there is a third party issue that we can…

us-east-2 goes down far, far less frequently than us-east-1. AWS doesn’t publicly release the outage numbers (they hold them very close to the chest) but some people have compiled the stats on their own if you poke around.

The regions provide the same functionality, so I see genuinely no downside or additional work to picking the 2 regions over the 2 regions.

It seems like one of those no brainer decisions to me. I take pride in being up when everyone else is down. 5 9s or bust, baby!

Re: Is Northern Virginia still the least reliable AWS region?

#70

Earlier quoted context omitted.

Bizarre way of making decisions. us-east-2 is objectively a better region to pick if you want US east, yet you feel safer picking use1 because “I’m safer making a worse decision that everyone understands is worse, as long as everyone else does it as well.”

It's about risk profile. The question isn't "which region goes down the least" but "how often will I be blamed for an outage." If you never get blamed for a US east outage, that's better than us-east-2 if that could get you blamed 0.5% of the time when it goes down and us1 isn't down or etc

But ise1 is down 4x more than use2 (AWS closely guards the numbers and won’t release them, but that is what I’ve seen from 3rd party analysis). Don’t you want your customers to say, “wow, half the internet was down today but XYZ service was up with no issues! I love them.”

I can’t tell if it’s you thinking this way, or if your company is setup to incentivize this. But either way, I think it’s suboptimal.

That’s not about “risk profile” of the business or making the right decision for the customer, that’s about risk profile of saving your own tail in the organizational gamesmanship sense. Which is a shame, tbh. For both the customer and for people making tech decisions.

I fully appreciate that some companies may encourage this behavior, and we all need a job so we have to work somewhere, but this type of thinking objectively leads to worse technology decisions and I hope I never have to work for a company that encourages this.

Edit: addressing blame when things go wrong… don’t you think it would be a better story to tell your boss that you did the right thing for the customer, rather than “I did this because everyone else does it, even though most of us agree it’s worse for the customer in general”. I would assume I’d get more blame for the 2nd decision than the 1st.

Post reply on HN