Live data from Hacker News

AWS North Virginia data center outage – resolved

cnbc.com

201–210 of 214 posts

Re: AWS North Virginia data center outage – resolved

#201
post #192

Earlier quoted context omitted.

Paper mills need a lot of heat energy to run the processes. Data centres produce a lot of heat. Sounds like a good combination? Cold water -> data centre cooling loop - > warm water -> paper mill with heat pumps to transform low-grade heat into the required temperatures -> profit

I can't believe any town would vote for a paper mill. It smells like a paper mill.

Well, they do provide jobs and tend to stay around for a while given the large investments needed to establish one. There is a big one about 15 km from where I live, it smells about the same as a wastewater treatment plant. There was one close to where I lived while in university as well (another era, another country) which mostly smelled of warm paper, no bad smells. All in all there are worse industries to have around.

Re: AWS North Virginia data center outage – resolved

#202
post #34

AWS’s US-East 1 continues to be the Achilles heel of the Internet. And while yes building across multiple regions and AZs is a thing, AWS has had a string of issues where US-East 1 has broader impacts, which makes things far less redundant and resilient than AWS implies.

> building across multiple regions and AZs is a thing If you do this for resiliency, be prepared to pay the capacity tax (2 regions means 2x capacity, 3 regions means 1.5x), have the machines already running in a multi-region setup (don't expect to be able to spin up instances or even get capacity during an outage), and ready to deal with the added complexity of multi-region hosting.

There’s all kinds of fun pitfalls with multi-AZ. Like you can create RDS subnets across multiple AZs but then you can’t remove an AZ. Which really sucks when your core database covers all 5 us-east-1 AZs and randomly can’t failover because you picked an instance type that use1-az4 can’t host.

Re: AWS North Virginia data center outage – resolved

#203

Earlier quoted context omitted.

Yes, you're right, but in my experience the boundary between the data plane and the control plane is not always clear, and especially unclear on these foundational and basic services. There were enough "surprisingly control-plane" IAM operations in the AWS services that I dealt with, so we had to exercise extreme caution during outages.

It's literally documented. Try reading it and educating yourself.

I worked there.

Even if I were the stupidest and least curious engineer around (and I was far from it), that's basically irrelevant to what you're scolding me for here…

As part of a team with both software development and operational responsibilities, like most teams at AWS, I had to deal not only with the consequences of my own imperfect knowledge, but also with the imperfect knowledge of my coworkers past and present.

Re: AWS North Virginia data center outage – resolved

#205

These things are dangerous. Someone who can take AWS down such as an employee can place a bet. These bets aren’t as innocent as they seem because the bettors can often influence or change the outcome.

> These things are dangerous. Someone who can take AWS down such as an employee can place a bet. Imagine if the betting website itself shuts down because AWS is down. (half joking I suppose though) > These bets aren’t as innocent as they seem because the bettors can often influence or change the outcome. Overall I agree with your statement that these betting markets also are able to incentivize a lot of insider tradi…

This is my retirement plan.

Get a job where I can affect something significant, mortgage everything I own to bet on it, then break it (get fired) and take the money and run.

Re: AWS North Virginia data center outage – resolved

#206

Earlier quoted context omitted.

This is almost definitely an issue of equipment failure. Cooling in datacenters is like everything else both over and under provisioned. It's overprovisioned in the sense that the big heat exchange units are N+1 (or in very critical and smaller load facilities 2N/3N). This is done because you need to regularly take these down for maintenance work and they have a relatively high failure rate compared to traditional DC…

I could totally get into “Ops Thriller” genre of novels like this.

It’s old but “The Cuckoo's Egg” was a great read and had a lot of this. Oh and it was true.

Re: AWS North Virginia data center outage – resolved

#207

Earlier quoted context omitted.

there are dozens of us!

There are often little bits of Neal Stephenson or Andy Weir novels which sound a little like this, describing a technical fault in a plot-driven way (often as a cascade), and I do find those to be uniquely enjoyable. I'm sure there are other authors who do similar things, though maybe "cloud/AI data center" stories should be its own micro-genre, given how crucial these things are to society.

Magic version of this: https://scp-wiki.wikidot.com/scp-5243

Re: AWS North Virginia data center outage – resolved

#208

Earlier quoted context omitted.

The cooling units dont fail just because they get to 100% duty cycle. That's pretty much "normal operation", you just get... higher efficiency coz the cooling side is warmer

Of course not. They fail above 100%. Some fail below 100% too.

That's not how anything works. Please stop spewing bollocks.
Post reply on HN