Live data from Hacker News

AWS S3 Outage

news.ycombinator.com

131–140 of 142 posts

Re: AWS S3 Outage

#132

Earlier quoted context omitted.

Amazon isn't a couple of guys in a garage. They have hordes of administrative personnel who could be tasked to update a status page. Companies 1/1000th the size of Amazon can manage it.

Yes and no. Sure, Amazon has administrative personnel. Sure, some of those administrative personnel would probably be happy to get paid extra to carry a pager and be summoned to work at 3AM to update a status page. But the last thing you want to do is put inaccurate information onto a status page; so mere administrative personnel isn't enough -- you'd need people who understand enough about the system to be able to w…

This is a situation where "we pay you, now do as you're told" comes in handy.

Not every job can be full of self-directed aspirational spiritual awakenings. If that were the case, nobody would deliver my dinner on a bike when it's -20ºF outside.

Re: AWS S3 Outage

#133

Earlier quoted context omitted.

Amazon isn't a couple of guys in a garage. They have hordes of administrative personnel who could be tasked to update a status page. Companies 1/1000th the size of Amazon can manage it.

Yes and no. Sure, Amazon has administrative personnel. Sure, some of those administrative personnel would probably be happy to get paid extra to carry a pager and be summoned to work at 3AM to update a status page. But the last thing you want to do is put inaccurate information onto a status page; so mere administrative personnel isn't enough -- you'd need people who understand enough about the system to be able to w…

>I'm guessing that the intersection of "administrative personnel", "willing to carry pagers" and "understand the internals of AWS services" is a very small set

Being a non-engineer doesn't mean they don't know anything about the technology. And they don't need to know the internals, just enough to convey information from the engineers managers to the public.

Plenty of other organizations manage resolving issues while transmitting information about the issue to other stakeholders.

Also, most administrative personnel have far less job opportunities than engineers. If they can get the engineers to carry pagers they can get a PR minion to carry one.

Re: AWS S3 Outage

#134

Earlier quoted context omitted.

What's stopping you from writing it?

Time and priorities. There is a difference between "I wish this thing existed, so I can use it/contribute to it" and "I need it so badly that I'm willing to spend a lot of time to make it production ready". So far S3 seems to be reliable enough...

Might be something to integrate into Libcloud [1] instead of rolling your own.

[1] https://libcloud.apache.org/

Re: AWS S3 Outage

#135
post #92

Seems Hasicorp is maybe affected by this as well. $ vagrant up Bringing machine '...' up with 'virtualbox' provider... ==> ...: Box 'debian/jessie64' could not be found. ... ...: Downloading: https://atlas.hashicorp.com/debian/boxes/jessie64/versions/8.1.0/providers/virtualbox.box An error occurred while downloading the remote file. The error message, if any, is reproduced below. Please fix this error and try again.…

VagrantCloud, now known as Atlas, is merely a redirector service for Vagrant box management. I host all my own boxes (on S3), but still use them for an easy means of sharing without having to remember a long-ass URL.

Re: AWS S3 Outage

#136

Earlier quoted context omitted.

Amazon isn't a couple of guys in a garage. They have hordes of administrative personnel who could be tasked to update a status page. Companies 1/1000th the size of Amazon can manage it.

Yes and no. Sure, Amazon has administrative personnel. Sure, some of those administrative personnel would probably be happy to get paid extra to carry a pager and be summoned to work at 3AM to update a status page. But the last thing you want to do is put inaccurate information onto a status page; so mere administrative personnel isn't enough -- you'd need people who understand enough about the system to be able to w…

Except "willing to carry pagers" is currently the basis of employment at Amazon, and not just for AWS but for whole chunks of their technical business. It's one of the many reasons why they have a pretty dire reputation (see plenty of discussions on here from former Amazon employees).

They also claim to have "customer obsession" as a leadership principle, this whole thread is an excellent example of that being failed in a big way.

Re: AWS S3 Outage

#137

Earlier quoted context omitted.

Amazon isn't a couple of guys in a garage. They have hordes of administrative personnel who could be tasked to update a status page. Companies 1/1000th the size of Amazon can manage it.

Yes and no. Sure, Amazon has administrative personnel. Sure, some of those administrative personnel would probably be happy to get paid extra to carry a pager and be summoned to work at 3AM to update a status page. But the last thing you want to do is put inaccurate information onto a status page; so mere administrative personnel isn't enough -- you'd need people who understand enough about the system to be able to w…

[deleted]

Re: AWS S3 Outage

#138
post #90

Earlier quoted context omitted.

Last one was 10 days ago. https://news.ycombinator.com/item?id=9980222

+1, Also it was down on leap second. http://mashable.com/2015/06/30/aws-disruption

To be fair, that was not really AWS's fault, nor (apparently) a leap second issue:

http://www.bgpexpert.com/article.php?article=167

https://twitter.com/Axcelx/status/616058414746202113

Re: AWS S3 Outage

#139

Earlier quoted context omitted.

I'm sure "elevated error rates" is the first alarm which goes off. And once they've put that description onto the status page, they're probably more worried about getting it fixed than going back and changing the wording.

Amazon isn't a couple of guys in a garage. They have hordes of administrative personnel who could be tasked to update a status page. Companies 1/1000th the size of Amazon can manage it.

You would be surprised to learn just how few people run your favorite web service.

Re: AWS S3 Outage

#140
post #46
post #19

Earlier quoted context omitted.

Ya, it bothers me that their status messages for major outages are simply "elevated error rates".

Whats frustrating is when you have customers who are also down because of the outage - but when you say Amazon is experiencing severe outages causing 50% of our requests to be dropped and there's not much we can do, it makes us look pretty bad when they they go to the amazon dashboard and only see "Elevated Error Rates."

In these cases I suggest saying "we are being severely affected by an Amazon outage"
Post reply on HN