Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

761–770 of 1001 posts

Re: AWS us-east-1 outage

#761
post #718
post #694

Earlier quoted context omitted.

>It seems bad now but I wonder how much worse it might be when no one actually has access to money because all financial traffic is going through AWS and it goes down. Most financial institutions are implementing their own clouds, I can't think of any major one that is reliant on public cloud to the extent transactions would stop. >Why hasn't the industry come up with an alternative? You mean like building datacenter…

> Most financial institutions are implementing their own clouds https://www.nasdaq.com/Nasdaq-AWS-cloud-announcement

That doesn't mean what you think it means.

The agreement is more of a hybrid cloud arrangement with AWS Outposts.

FTA:

>Core to Nasdaq’s move to AWS will be AWS Outposts, which extend AWS infrastructure, services, APIs, and tools to virtually any datacenter, co-location space, or on-premises facility. Nasdaq plans to incorporate AWS Outposts directly into its core network to deliver ultra-low-latency edge compute capabilities from its primary data center in Carteret, NJ.

They are also starting small, with Nasdaq MRX

This is much less about moving NASDAQ (or other exchanges) to be fully owned/maintained by Amazon, and more about wanting to take advantage of development tooling and resources and services AWS provides, but within the confines of an owned/maintained data center. I'm sure as this partnership grows, racks and racks will be in Amazon's data centers too, but this is a hybrid approach.

I would also bet a significant amount of money that when NASDAQ does go full "cloud" (or hybrid, as it were), it won't be in the same US-east region co-mingling with the rest of the consumer web, but with its own redundant services and connections and networking stack.

NASDAQ wants to modernize its infrastructure but it absolutely doesn't want to offload it to a cloud provider. That's why it's a hybrid partnership.

Re: AWS us-east-1 outage

#762

I think now is a good time to reiterate the danger of companies just throwing all of their operational resilience and sustainability over the wall and trusting someone else with their entire existence. It's wild to me that so many high performing businesses simply don't have a plan for when the cloud goes down. Some of my contacts are telling me that these outages have teams of thousands of people completely prevente…

We're too busy working generating our own electricity and designing our own CPUs.

Re: AWS us-east-1 outage

#763

I think now is a good time to reiterate the danger of companies just throwing all of their operational resilience and sustainability over the wall and trusting someone else with their entire existence. It's wild to me that so many high performing businesses simply don't have a plan for when the cloud goes down. Some of my contacts are telling me that these outages have teams of thousands of people completely prevente…

This appears to be a single region outage - us-east-1. AWS supports as much redundancy as you want. You can be redundant between multiple Availability Zones in a single Region or you can be redundant among 1, 2 or even 25 regions throughout the world.

Multiple-region redundancy costs more both in initial planning/setup as well as monthly fees so a lot of AWS customers choose to just not do it.

Re: AWS us-east-1 outage

#764

I think now is a good time to reiterate the danger of companies just throwing all of their operational resilience and sustainability over the wall and trusting someone else with their entire existence. It's wild to me that so many high performing businesses simply don't have a plan for when the cloud goes down. Some of my contacts are telling me that these outages have teams of thousands of people completely prevente…

In my opinion there is a lack of talent in these industries for building out there own resilient systems. IT people and engineers get lazy.

> IT people and engineers get lazy.

Companies do not change their whole strategy from a capex-driven traditional self-hosting environment to opex-driven cloud hosting because their IT people are lazy; it is typically an exec-level decision.

Re: AWS us-east-1 outage

#765

I think now is a good time to reiterate the danger of companies just throwing all of their operational resilience and sustainability over the wall and trusting someone else with their entire existence. It's wild to me that so many high performing businesses simply don't have a plan for when the cloud goes down. Some of my contacts are telling me that these outages have teams of thousands of people completely prevente…

This seems like an insane stance to have, it's like saying businesses should ship their own stock, using their own drivers, and their in-house made cars and planes and in-house trained pilots.

Heck, why stop at having servers on-site? Cast your own silicon waffers, after all you don't want spectrum exploits.

Because you are worst at it. If a specialist is this bad, and the market is fully open, then it's because the problem is hard.

AWS has fewer outages in one zone alone than the best self-hosted institutions, your facebooks and petagons. In-house servers would lead to an insane amount of outage.

And guess what? AWS (and all other IAAS providers) will beg you to use multiple region because of this. The team/person that has millions of dollars a day staked on a single AWS region is an idiot and could not be entrusted to order a gaming PC from newegg, let alone run an in-house datacenter.

edit: I will add that AWS specifically is meh and I wouldn't use it myself, there's better IASS. But it's insanity to even imagine self-hosted is more reliable than using even the shittiest of IASS providers.

Re: AWS us-east-1 outage

#766

Earlier quoted context omitted.

They are still lying about it, the issues are not only affecting the console but also AWS operations such as S3 puts. S3 still shows green.

Yep, I am seeing failures on IAM as well: aws iam list-policies An error occurred (503) when calling the ListPolicies operation (reached max retries: 2): Service Unavailable

Same here. Kubernetes pods running in EKS are (intermittently) failing to get IAM credentials via the ServiceAccount integration.

Re: AWS us-east-1 outage

#767

Earlier quoted context omitted.

Cynical/realist take: Take responsibility and then hope your bosses already love you, you can immediately both come with a way to prevent it from happening again, and convince them to give you the resources to implement it. Otherwise your responsibility is, unfortunately, just blood in the water for someone else to do all of that to protect the company against you and springboard their reputation on the descent of yo…

This seems like an absolutely horrid way of working or doing 'office politics'.

Yes, and I personally have worked in environments that do just that. They said they didn't, but with management "personalities" plus stack ranking, you know damn well that they did.

Re: AWS us-east-1 outage

#768

Earlier quoted context omitted.

It's popular to upvote this during outages, because it fits a narrative. The truth (as always) is more complex: * No, this isn't the broad culture. It's not even a blip. These are EXCEPTIONAL circumstances by extremely bad teams that - if and when found out - would be intervened dramatically. * The broad culture is blameless post-mortems. Not whose fault is it. But what was the problem and how to fix it. And one of t…

Well, the narrative is sort of what Amazon is asking for, heh? The whole us-east-1 management console is gone, what is Amazon posting for the management console on their website? "Service degradation" It's not a degradation if it's outright down. Use the red status a little bit more often, this is a "disruption", not a "degradation".

Well you know, like when a rocket explode, it's a sudden and "unexpected rapid disassembly" or something...

And a cleaner is called a "floor technician".

Nothing really out of the ordinary for a service to be called degraded while "hey, the cache might still be working right?" ... or "Well you know, it works every other day except today, so it's just degradation" :-)

Re: AWS us-east-1 outage

#769
post #245

Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…

Uh, four minutes to identify the root cause? Damn, those guys are on fire.

Outage started at 731 PST from our monitoring. They are on fire, but not in a good way.

Re: AWS us-east-1 outage

#770

I think now is a good time to reiterate the danger of companies just throwing all of their operational resilience and sustainability over the wall and trusting someone else with their entire existence. It's wild to me that so many high performing businesses simply don't have a plan for when the cloud goes down. Some of my contacts are telling me that these outages have teams of thousands of people completely prevente…

> Why hasn't the industry come up with an alternative?

The cloud is the solution to self managed data centers. Their value proposition is appealing: Focus on your core business and let us handle infrastructure for you.

This fits the needs of most small and medium sized businesses, there's no reason not to use the cloud and spend time and money on building and operating private data centers when the (perceived) chances of outages are so small.

Then, companies grow to a certain size where the benefits of having a self managed data center begins to outweight not having one. But at this point this becomes more of a strategic/political decision than merely a technical one, so it's not an easy shift.

Post reply on HN