Could someone explain to me why they don't build these things near oceans? Like nuclear plants that need plenty cooling capacity too Two loop cycle with heat exchanger to get rid of the heat
AWS North Virginia data center outage – resolved
21–30 of 214 posts
Re: AWS North Virginia data center outage – resolved
#22Re: AWS North Virginia data center outage – resolved
#23Re: AWS North Virginia data center outage – resolved
#24I thought cooling was pretty much pre-planned in any data center, and you simply don't install more stuff than you can cool? So did some cooling equipment fail here or was there an external reason for the overheating? Or does Amazon overbook the cooling in their data centers?
This is almost definitely an issue of equipment failure. Cooling in datacenters is like everything else both over and under provisioned. It's overprovisioned in the sense that the big heat exchange units are N+1 (or in very critical and smaller load facilities 2N/3N). This is done because you need to regularly take these down for maintenance work and they have a relatively high failure rate compared to traditional DC…
Re: AWS North Virginia data center outage – resolved
#25Earlier quoted context omitted.
This is almost definitely an issue of equipment failure. Cooling in datacenters is like everything else both over and under provisioned. It's overprovisioned in the sense that the big heat exchange units are N+1 (or in very critical and smaller load facilities 2N/3N). This is done because you need to regularly take these down for maintenance work and they have a relatively high failure rate compared to traditional DC…
I'd expect someone like AWS to just throttle machines before overloading their cooling. Because they probably can do that, while e.g. a data center that just rents the space can't really throttle their customers nicely.
But they did load-shed. Perhaps not soon enough, but the reason this is publicly known is because they reduced the amount of heat being produced.
Re: AWS North Virginia data center outage – resolved
#26Earlier quoted context omitted.
This is almost definitely an issue of equipment failure. Cooling in datacenters is like everything else both over and under provisioned. It's overprovisioned in the sense that the big heat exchange units are N+1 (or in very critical and smaller load facilities 2N/3N). This is done because you need to regularly take these down for maintenance work and they have a relatively high failure rate compared to traditional DC…
Shouldn't there be a feedback system here preventing the scheduling of loads when cooling is degraded?
But this is the physical world, shit happens.
The algorithm didn't know that fuse was lose and fine at 50% duty cycle but was high resistance and going to blow at 100%.
Re: AWS North Virginia data center outage – resolved
#27Could someone explain to me why they don't build these things near oceans? Like nuclear plants that need plenty cooling capacity too Two loop cycle with heat exchanger to get rid of the heat
Amusingly I've been part of two critical downtime heating incidents at two different datacenters: one was when Hosting.com's SOMA datacenter got so hot that they were using hoses on the roof to cool it down; and the second one was when Alibaba's Chai Wan datacenter got so hot everything running there went down, including the control plane. So I imagine the proximity to the ocean does not yield any additional advantag…
Re: AWS North Virginia data center outage – resolved
#28Earlier quoted context omitted.
This is almost definitely an issue of equipment failure. Cooling in datacenters is like everything else both over and under provisioned. It's overprovisioned in the sense that the big heat exchange units are N+1 (or in very critical and smaller load facilities 2N/3N). This is done because you need to regularly take these down for maintenance work and they have a relatively high failure rate compared to traditional DC…
I'd expect someone like AWS to just throttle machines before overloading their cooling. Because they probably can do that, while e.g. a data center that just rents the space can't really throttle their customers nicely.
A decade ago it was trivial to just tell the hypervisor to reduce the cpu fraction of all VMs by half and leave half unallocated. Now, it's much more complicated and definitely would be user visible.
Re: AWS North Virginia data center outage – resolved
#29Earlier quoted context omitted.
This is almost definitely an issue of equipment failure. Cooling in datacenters is like everything else both over and under provisioned. It's overprovisioned in the sense that the big heat exchange units are N+1 (or in very critical and smaller load facilities 2N/3N). This is done because you need to regularly take these down for maintenance work and they have a relatively high failure rate compared to traditional DC…
The cooling units dont fail just because they get to 100% duty cycle. That's pretty much "normal operation", you just get... higher efficiency coz the cooling side is warmer
Some fail below 100% too.
Re: AWS North Virginia data center outage – resolved
#30Could someone explain to me why they don't build these things near oceans? Like nuclear plants that need plenty cooling capacity too Two loop cycle with heat exchanger to get rid of the heat