Live data from Hacker News

Summary of the AWS Service Event in the Sydney Region

aws.amazon.com

21–30 of 41 posts

Re: Summary of the AWS Service Event in the Sydney Region

#21

Why, oh why do they report times in PDT rather than AEST (the zone of the affected area) or UTC (the standard everything else is based on)? (Mutter, mutter, … something about Americans and their timezones … and northern hemispherians and their seasons …)

Because this post-mortem was written by an American in Seattle

Re: Summary of the AWS Service Event in the Sydney Region

#22
post #15

Our instances with ap-southeast-2 were out for around 12 hours. We used multiple availability zones and it didn't prevent downtime at all. It's very interesting the difference between AWS and Google outage responses. AWS is down for 12+ hours for some customers, force each customer to chase service level credits and sign off the postmortem with a nameless & faceless "-The AWS Team". Not one person at AWS was willing…

people vote with their dollars and their votes are overwhelmingly telling amazon they're doing a great job.

Re: Summary of the AWS Service Event in the Sydney Region

#23
post #15

Our instances with ap-southeast-2 were out for around 12 hours. We used multiple availability zones and it didn't prevent downtime at all. It's very interesting the difference between AWS and Google outage responses. AWS is down for 12+ hours for some customers, force each customer to chase service level credits and sign off the postmortem with a nameless & faceless "-The AWS Team". Not one person at AWS was willing…

> We used multiple availability zones and it didn't prevent downtime at all.

Can you explain this a little more? Amazon says this only affected one AZ, and they specifically note:

    For this event, customers that were running their applications across multiple 
    Availability Zones in the Region were able to maintain availability throughout 
    the event.

Re: Summary of the AWS Service Event in the Sydney Region

#24
post #15

Our instances with ap-southeast-2 were out for around 12 hours. We used multiple availability zones and it didn't prevent downtime at all. It's very interesting the difference between AWS and Google outage responses. AWS is down for 12+ hours for some customers, force each customer to chase service level credits and sign off the postmortem with a nameless & faceless "-The AWS Team". Not one person at AWS was willing…

I wonder if due to the scale of AWS and certain AWS customers, if AWS signs a post mortem with -Fred would a very large AWS customer have the pull to say "Amazon, fire Fred"? Just curious if anonymous postmortems are company policy at certain places and why that might be.

[deleted]

Re: Summary of the AWS Service Event in the Sydney Region

#25
post #15

Our instances with ap-southeast-2 were out for around 12 hours. We used multiple availability zones and it didn't prevent downtime at all. It's very interesting the difference between AWS and Google outage responses. AWS is down for 12+ hours for some customers, force each customer to chase service level credits and sign off the postmortem with a nameless & faceless "-The AWS Team". Not one person at AWS was willing…

[deleted]

Re: Summary of the AWS Service Event in the Sydney Region

#26

>> The specific signature of this weekend’s utility power failure resulted in an unusually long voltage sag (rather than a complete outage) It is false to assume that the state of the electrical supply is either on or off. This may come as a surprise, but not to me. In 2008, Eskom (South Africa's electricity suppliers) experienced similar faults. The mains supply voltage is 220v here. At one point, some devices start…

It doesn't sound like they were treating it as binary, they have breakers in place for brownout detection too -- those breakers just weren't triggered fast/early enough in the brown out.

Re: Summary of the AWS Service Event in the Sydney Region

#27
post #15

Our instances with ap-southeast-2 were out for around 12 hours. We used multiple availability zones and it didn't prevent downtime at all. It's very interesting the difference between AWS and Google outage responses. AWS is down for 12+ hours for some customers, force each customer to chase service level credits and sign off the postmortem with a nameless & faceless "-The AWS Team". Not one person at AWS was willing…

I independently monitor availability of 150 public cloud services and only observed 1.73 hours downtime for this event. This is the first EC2 outage I've observed in any region for over 6 months. According to my stats, in 2015 EC2 was highly available with 6 of 9 regions (including every US and EU region) having no outages, and total average service availability of 99.998% (78% of downtime in sa-east-1). I haven't observed a single outage in us-west-2 or eu-west-1 in nearly 3 years. This compared to 16-33 minutes of downtime in every region for GCE in 2015, and total average availability of 99.995%. Additionally, since 2013 I've never observed a global EC2 outage like the 4/11/2016 GCE event.

https://cloudharmony.com/status-for-aws

Re: Summary of the AWS Service Event in the Sydney Region

#28
post #19

I love reading about problems like these, it's great that Amazon is forthcoming about them. There's always some new wrinkle. E.g. in this case, in normal operation, power from the utility power grid spins a flywheel. When the grid fails, the flywheel provides a holdover until Amazon's diesel generators can start. But in this failure the voltage from the grid sagged, rather than going away completely. The breaker isol…

Something else must have failed or was not properly configured, because backup Diesel generators should kick-in after 2-15 seconds of voltage drop, regardless of the flywheel. The flywheel is used in critical systems to cover only that <1m gap.

Amazon addresses that in their report. Each UPS is fed by generator power and grid power. Because the UPSes had been forced to try to supply the grid, and because they are giant spinning weights that you really don't want to go wrong and kill someone or destroy property, they did a safety inspection before powering them back up, which meant a delay before the facility could be supplied by generator power.

Re: Summary of the AWS Service Event in the Sydney Region

#29
post #27
post #15

Our instances with ap-southeast-2 were out for around 12 hours. We used multiple availability zones and it didn't prevent downtime at all. It's very interesting the difference between AWS and Google outage responses. AWS is down for 12+ hours for some customers, force each customer to chase service level credits and sign off the postmortem with a nameless & faceless "-The AWS Team". Not one person at AWS was willing…

I independently monitor availability of 150 public cloud services and only observed 1.73 hours downtime for this event. This is the first EC2 outage I've observed in any region for over 6 months. According to my stats, in 2015 EC2 was highly available with 6 of 9 regions (including every US and EU region) having no outages, and total average service availability of 99.998% (78% of downtime in sa-east-1). I haven't ob…

Most of the criticism seemed to be centered around communication/PR so statistics don't really address that. Definitely important when considering a provider though.

While the numbers are nice, if some outages only impact a subset of customers and your monitoring accounts aren't one of them it's hard to determine how good your monitoring data really is. If he was impacted by a 12 hour outage and you only show ~2 hours that's a really significant difference.

I guess it really depends on how you monitor and how comprehensive it is. Do you monitor from multiple ISPs on different network paths in multiple regions/countries? Do you monitor each of the services under different load conditions and monitor multiple accounts? Sometimes "up" only tells part of the story.

Re: Summary of the AWS Service Event in the Sydney Region

#30
post #15

Our instances with ap-southeast-2 were out for around 12 hours. We used multiple availability zones and it didn't prevent downtime at all. It's very interesting the difference between AWS and Google outage responses. AWS is down for 12+ hours for some customers, force each customer to chase service level credits and sign off the postmortem with a nameless & faceless "-The AWS Team". Not one person at AWS was willing…

people vote with their dollars and their votes are overwhelmingly telling amazon they're doing a great job.

For the record I agree with you but existing customers who are heavily invested in AWS would find it difficult to vote with their dollars.
Post reply on HN