Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

751–760 of 1001 posts

Re: AWS us-east-1 outage

#751
post #377

Earlier quoted context omitted.

Would they? Having 3 outages in a year sounds like an organization problem. Not enough safeguards to prevent very routine human errors. But instead of worrying about that we just assign a guy to take the fall

Well if John caused 3 outages and and his peers Sally and Mike each caused 0, it's worth taking a deeper look. There's a real possibility he's getting screwed by a messed up org, also he could be doing slapdash work or he seriously might not undertsand the seriousness of an outage.

Worth a look, certainly. Also very possible that this John is upfront about honest postmortems and like a good leader takes the blame, whereas Sally and Mike are out all day playing politics looking for how to shift blame so nothing has their name attached. Most larger companies that's how it goes.

Re: AWS us-east-1 outage

#752
post #214

Earlier quoted context omitted.

If you're not multi-cloud in 2021 and are expecting 5-9's, I feel bad for you.

If you're having SLA problems I feel bad for you son I got two 9 problems cuz of us-east-1

> ~~I got two 9 problems cuz of us-east-1~~

I left my two nines problems in us-east-1

Re: AWS us-east-1 outage

#753
post #308

Earlier quoted context omitted.

Haha... This bring back memories. It really depends on the org. I've had push backs on my postmortems before because of phrasing that could be constituted as laying some of the blame on some person/team when it's supposed to be blameless. And for a long time, it was fairly blameless. You would still be punished with the extra work of writing high quality postmortems, but I have seen people accidentally bring down cri…

We're in a situation where the balls of mud made people afraid to touch some things in the system. As experiences and processes have improved we've started to crack back into those things and guess what, when you are being groomed to own a process you're going to fuck it up from time to time. Objectively, we're still breaking production less often per year than other teams, but we are breaking it, and that's novel be…

Or problems just persisting, because the fix is easy, but explaining it to others who do not work on the system are hard. Esp. justifying why it won't cause an issue, and being told that the fixes need to be done via scripts that will only ever be used once, but nevertheless needs to be code reviewed and tested...

I wanted to be proactive and fix things before they became an issue, but such things just drained life out of me, to the point I just left.

Re: AWS us-east-1 outage

#754

I think now is a good time to reiterate the danger of companies just throwing all of their operational resilience and sustainability over the wall and trusting someone else with their entire existence. It's wild to me that so many high performing businesses simply don't have a plan for when the cloud goes down. Some of my contacts are telling me that these outages have teams of thousands of people completely prevente…

Because your own datacenters cant go down?

Re: AWS us-east-1 outage

#755

I think now is a good time to reiterate the danger of companies just throwing all of their operational resilience and sustainability over the wall and trusting someone else with their entire existence. It's wild to me that so many high performing businesses simply don't have a plan for when the cloud goes down. Some of my contacts are telling me that these outages have teams of thousands of people completely prevente…

Because the expected value of using AWS is greater than the expected value of self-hosting. It's not that nobody's ever heard of running on their own metal. Look back at what everyone did before AWS, and how fast they ran screaming away from it as soon as they could. Once you didn't have to do that any more, it's just so much better that the rare outages are worth it for the vast majority of startups.

Medical devices, banks, the military, etc. should generally run on their own hardware. The next photo-sharing app? It's just not worth it until they hit tremendous scale.

Re: AWS us-east-1 outage

#757

I think now is a good time to reiterate the danger of companies just throwing all of their operational resilience and sustainability over the wall and trusting someone else with their entire existence. It's wild to me that so many high performing businesses simply don't have a plan for when the cloud goes down. Some of my contacts are telling me that these outages have teams of thousands of people completely prevente…

Many of those businesses wouldn’t have existed in the first place without simplicity offered by cloud.

Re: AWS us-east-1 outage

#758

I think now is a good time to reiterate the danger of companies just throwing all of their operational resilience and sustainability over the wall and trusting someone else with their entire existence. It's wild to me that so many high performing businesses simply don't have a plan for when the cloud goes down. Some of my contacts are telling me that these outages have teams of thousands of people completely prevente…

> tens of million dollars of profit are simply vanishing

vanishing or delayed six hours? I mean

Re: AWS us-east-1 outage

#759

I think now is a good time to reiterate the danger of companies just throwing all of their operational resilience and sustainability over the wall and trusting someone else with their entire existence. It's wild to me that so many high performing businesses simply don't have a plan for when the cloud goes down. Some of my contacts are telling me that these outages have teams of thousands of people completely prevente…

So you're saying companies should start moving their infrastructure to the blockchain?

Re: AWS us-east-1 outage

#760

I think now is a good time to reiterate the danger of companies just throwing all of their operational resilience and sustainability over the wall and trusting someone else with their entire existence. It's wild to me that so many high performing businesses simply don't have a plan for when the cloud goes down. Some of my contacts are telling me that these outages have teams of thousands of people completely prevente…

In my opinion there is a lack of talent in these industries for building out there own resilient systems. IT people and engineers get lazy.

No lazier than anyone else, there's just not enough of us, in general and per company.
Post reply on HN