Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

581–590 of 1001 posts

Re: AWS us-east-1 outage

#581

noob question: Aren't companies using several regions for availability and redundancy?

I'm seeing outages across several regions for certain services (SNS), so cross-region failover doesn't necessarily help here.

Additionally, for complex apps, automatic cross-region disaster recovery can take tens or even hundreds of dev years, something most small to midsized companies can't afford.

Re: AWS us-east-1 outage

#582
post #378

Earlier quoted context omitted.

Would they? Having 3 outages in a year sounds like an organization problem. Not enough safeguards to prevent very routine human errors. But instead of worrying about that we just assign a guy to take the fall

If you work in a technical role and you _don't_ have the ability to break something, you're unlikely to be contributing in a significant way. Likely that would make you a junior developer whose every line of code is heavily scrutinized. Engineers should be experts and you should be able to trust them to make reasonable choices about the management of their projects. That doesn't mean there can't be some checks in pla…

> Which one provides more value to an organization?

Neither, they both provide the same value in the long term.

Senior engineers cannot execute on everything they commit to without having a team of engineers they work with. If nobody trains junior engineers, the discipline would go extinct.

Senior engineers provide value by building guardrails to enable junior engineers to provide value by delivering with more confidence.

Re: AWS us-east-1 outage

#583
It's funny but when I saw "AWS Outage" breaking, my first thought was "I bet it's US-east-1 again."

I know it's cheap but seriously... not worth it. Many of us have the scars to prove this.

Re: AWS us-east-1 outage

#584

Does anyone know why Google is showing the same spike on down detector as everything else? How does Google depend on AWS? https://downdetector.com/status/google/

It’s because Down Detector works off of user reports rather than automatically detecting outages somehow. So, every time a major service goes down (whether infrastructure like AWS or Cloudflare, or user-facing like YouTube or Facebook), some users will blame Google, ISPs, cellular providers, or some other unrelated service.

Re: AWS us-east-1 outage

#585

Does anyone know why Google is showing the same spike on down detector as everything else? How does Google depend on AWS? https://downdetector.com/status/google/

Some google sheets functions aren't updating in a timely manner for me. Maybe people google as a backup for aws and they have to throttle certain service from a higher load

Re: AWS us-east-1 outage

#587
Well I got to bugger off home early so good job Amazon.

Edit: to be clear this is because I’m utterly helplessly unable to do anything at the moment.

Re: AWS us-east-1 outage

#588
post #245

Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…

> This issue is affecting the global console landing page, which is also hosted in US-EAST-1 Even this little tidbit is a bit of a wtf for me. Why do they consider it ok to have anything hosted in a single region? At a different (unnamed) FAANG, we considered it unacceptable to have anything depend on a single region. Even the dinky little volunteer-run thing which ran https://internal.site.example/~someEngineer was…

I just want to serve 5 terabytes of data

Re: AWS us-east-1 outage

#589

Earlier quoted context omitted.

I worked at Walmart Technology. I bravely wrote post mortem documents owning the fault of my team (100+ people), owning both technically and also culturally as their leader. I put together a plan to fix it and executed it. Thought that was the right thing to do. This happend two times in my 10 year career there. Both times I was called out as a failure in my performance eval. Second time, I resigned and told them to…

That's shockingly stupid. I also worked for a major Walmart IT services vendor in another life, and we always had to be careful about how we handled them, because they didn't always show a lot of respect for vendors. On another note, thanks for building some awesome stuff -- walmart.com is awesome. I have both Prime and whatever-they're-currently-calling Walmart's version and I love that Walmart doesn't appear to mix…

Is WalMart.com awesome?

Re: AWS us-east-1 outage

#590
post #230
post #38

Earlier quoted context omitted.

Those five 9s don't come easy. Sometimes you have to prop them up :)

It’s hard to measure what five-9 is because you have to wait around until a 0.00001 occurs. Incentivizing post-mortems are absolutely critical in this case.

It's 0.001; the first 2 9's count.

  5N  = 99.999%
  3N  = 99.9%
  1N5 = 95%
5N is <43m12s downtime per month.
Post reply on HN