Live data from Hacker News

AWS North Virginia data center outage – resolved

cnbc.com

41–50 of 214 posts

Re: AWS North Virginia data center outage – resolved

#41

Earlier quoted context omitted.

What happens when the backup breaks?

You have a back up for the back up backup. Turtles all the way down. At AWS scale even unlikely hardware events become more common I guess.

Each turtle gives them another 9. How many 9s are they down due to incidents over the past year?

Re: AWS North Virginia data center outage – resolved

#42
post #34

AWS’s US-East 1 continues to be the Achilles heel of the Internet. And while yes building across multiple regions and AZs is a thing, AWS has had a string of issues where US-East 1 has broader impacts, which makes things far less redundant and resilient than AWS implies.

anecdotally (well, more "second-hand-ly i heard that..." it sounds like there were some carry-on effects on us-east-2 as a result of people migrating over from us-east-1, so, yeah... kinda hilarious how the multiple region / AZ thing is just so plainly a façade, but yet we all seem to just collectively believe in it as an article of faith in the Cloud Religion... or whatever...

Re: AWS North Virginia data center outage – resolved

#43

These things are dangerous. Someone who can take AWS down such as an employee can place a bet. These bets aren’t as innocent as they seem because the bettors can often influence or change the outcome.

It's a good thing big tech hires for ethical engineers and not ones that only care about money or social status.

Re: AWS North Virginia data center outage – resolved

#44

These things are dangerous. Someone who can take AWS down such as an employee can place a bet. These bets aren’t as innocent as they seem because the bettors can often influence or change the outcome.

> These things are dangerous. Someone who can take AWS down such as an employee can place a bet.

Imagine if the betting website itself shuts down because AWS is down. (half joking I suppose though)

> These bets aren’t as innocent as they seem because the bettors can often influence or change the outcome.

Overall I agree with your statement that these betting markets also are able to incentivize a lot of insider trading and one can say negative scenarios as this has given them an incentive to capitalize on that.

Re: AWS North Virginia data center outage – resolved

#45
post #4

I thought cooling was pretty much pre-planned in any data center, and you simply don't install more stuff than you can cool? So did some cooling equipment fail here or was there an external reason for the overheating? Or does Amazon overbook the cooling in their data centers?

This is almost definitely an issue of equipment failure. Cooling in datacenters is like everything else both over and under provisioned. It's overprovisioned in the sense that the big heat exchange units are N+1 (or in very critical and smaller load facilities 2N/3N). This is done because you need to regularly take these down for maintenance work and they have a relatively high failure rate compared to traditional DC…

I would have thought with all the data centers being built the parts for cooling systems would be standardized with replacements available from Grainger immediately.

Re: AWS North Virginia data center outage – resolved

#49
Once known for having super reliable services, I've heard this company is scrambling to re hire some of the engineers they overconfidently "replaced" with AI.

When customers pay for cloud services, they expect them to be maintained by competent engineers.

edit: Not sure why the downvotes. If you fire the engineers that have been keeping your systems running reliably for years, what do you expect to happen?

Re: AWS North Virginia data center outage – resolved

#50

Earlier quoted context omitted.

One of the data center's cooling loops broke.

No backups?

I once worked at a company that had a wealth of backups. A backup generator, backup batteries as the generator takes a few seconds to start, a contract for emergency fuel deliveries, a complete failover data centre full of hot standby hardware, 24/7 ops presence, UPSes on the ops PCs just in case, weekly checks that the generators start, quarterly checks by turning off the breakers to the data centre, and so on.

It wasn't until a real incident that we learned: (a) the system wasn't resilient to the utility power going on-off-on-off-on-off as each 'off' drained the batteries while the generator started, and each 'on' made the generator shut down again; (b) the ops PCs were on UPSes but their monitors weren't (C13 vs C5 power connector) and (c) the generator couldn't be refuelled while running.

Even if you've got backup systems and you test them - you can never be 100% sure.

Post reply on HN