How many nines of are we at this year?
AWS North Virginia data center outage – resolved
61–70 of 214 posts
Re: AWS North Virginia data center outage – resolved
#62These things are dangerous. Someone who can take AWS down such as an employee can place a bet. These bets aren’t as innocent as they seem because the bettors can often influence or change the outcome.
It's a good thing big tech hires for ethical engineers and not ones that only care about money or social status.
Re: AWS North Virginia data center outage – resolved
#63Earlier quoted context omitted.
The cooling units dont fail just because they get to 100% duty cycle. That's pretty much "normal operation", you just get... higher efficiency coz the cooling side is warmer
Of course not. They fail above 100%. Some fail below 100% too.
Re: AWS North Virginia data center outage – resolved
#64Re: AWS North Virginia data center outage – resolved
#65It's always East 1... Jokes aside I don't understand how often east-1 is taken down compared to other regions. Like it should be pretty similar to other regions architecture wise.
Isn't east one the "core" datacenter and also the oldest? I'd imagine it has more load than the other regions and also has more tech debt and architectural / engineering debt because they had less experience when they built it. Also iirc some services rely on east-1 as a single point of failure for configuration (like IAM or some S3 stuff?)
Re: AWS North Virginia data center outage – resolved
#66Earlier quoted context omitted.
So Ashburn VA is a datacenter hub because the very first non-government Internet Exchange Point (IXP) anywhere in the world was there ( https://en.wikipedia.org/wiki/MAE-East ). Back in the 1990's something like half of all internet traffic all over the world hit MAE-East. That in turn made AWS put their first region there (us-east-1 preceded eu-west-1 by 2 years and us-west-1 by 3 years). Then because there were lot…
The underlying reason is more that by being in us east coast you have about equal latency for customers in us west coast and Europe. That's a very large population covered from a single site. If you're building a single datacenter site this is where you start building first.
Re: AWS North Virginia data center outage – resolved
#67Earlier quoted context omitted.
No backups?
I once worked at a company that had a wealth of backups. A backup generator, backup batteries as the generator takes a few seconds to start, a contract for emergency fuel deliveries, a complete failover data centre full of hot standby hardware, 24/7 ops presence, UPSes on the ops PCs just in case, weekly checks that the generators start, quarterly checks by turning off the breakers to the data centre, and so on. It w…
Re: AWS North Virginia data center outage – resolved
#68Re: AWS North Virginia data center outage – resolved
#69Earlier quoted context omitted.
This is almost definitely an issue of equipment failure. Cooling in datacenters is like everything else both over and under provisioned. It's overprovisioned in the sense that the big heat exchange units are N+1 (or in very critical and smaller load facilities 2N/3N). This is done because you need to regularly take these down for maintenance work and they have a relatively high failure rate compared to traditional DC…
This is a great writeup! thank you!! Reminds when i did noogler training back in the day and one of the talks described a cascading failure at a datacenter, starting with a cat which was too curious near a power conditioner, and briefly conducted
Its cold up here in the winter, sadly, the residual heat from even totally passive components like switch gear is enough to warm things up enough to attract them. .001% of 1MW of power is still quite warm. (I have no idea how much switchgear leaks but i know they are warm even in winter outdoors).
And, yeah, the rest of the writeup is also an amalgamation of some panic-inducing experiences in my life.