My favorite sentence: "Our networking clients have well tested request back-off behaviors that are designed to allow our systems to recover from these sorts of congestion events, but, a latent issue prevented these clients from adequately backing off during this event."
I saw pleeeeeenty of untested code at Amazon/AWS. Looking back it was almost like the most important services/code had the least amount of testing. While internal boondoggle projects (I worked on a couple) had complicated test plans and debates about coverage metrics.
Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
211–220 of 410 posts
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#212Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#213My company uses AWS. We had significant degradation for many of their APIs for over six hours, having a substantive impact on our business. The entire time their outage board was solid green. We were in touch with their support people and knew it was bad but were under NDA not to discuss it with anyone. Of course problems and outages are going to happen, but saying they have five nines (99.999) uptime as measured by…
> The entire time their outage board was solid green Unless you're talking about some board other than the Service Health Dashboard, this isn't true. They dropped EC2 down to degraded pretty early on. I bemusedly noted in our corporate Slack that every time I refreshed the SHD, another service was listed as degraded. Then they added the giant banner at the top. Their slight delay in updating the SHD at the beginning…
Sagemaker, for example, was down all day. I was dead in the water on a modeling project that required GPUs. It relied on EC2, but nobody there even thought to update the status? WTF. This is clearly executives incentivized to let a bug persist. This is because the bug is actually a feature for misleading customers and maximizing profits.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#214I’ve been running platform teams on aws now for 10 years, and working in aws for 13. For anyone looking for guidance on how to avoid this, here’s the advice I give startups I advise. First, if you can, avoid us-east-1. Yes, you’ll miss new features, but it’s also the least stable region. Second, go multi AZ for production workloads. Safety of your customer’s data is your ethical responsibility. Protect it, back it up…
A wait for X provider to fix it for you situation is infinitely more stressful than an 'I have played myself, I will now take action' situation.
Situations out of your (immediate) resolution control feel infinitely worse, even if the customer impact in practice of your fault vs cloud fault is the same.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#215Earlier quoted context omitted.
>>It's smart politics -- I don't blame them Um, so you think straight-up lying is good politics? Any 7-year old knows that telling a lie when you broke something makes you look better superficially, especially if you get away with it. That does not mean that we should think it is a good idea to tell lies when you break things. It sure as hell isn't smart politics in my book. It is straight-up disqualifying to do busi…
You don’t know what you’re talking about. AWS spends a lot of time thinking about this problem in service to their customers. How do you reduce the status of millions of machines, the software they run, and the interconnected-ness of those systems to a single graphical indicator? It would be dumb and useless to turn something red every single time anything had a problem. Literally there are hundreds of things broken…
A good low-hanging fruit would be, when the outage is significant enough to have reached the media, you turn the dot red.
Dishonesty is what we're talking about here. Not the gradient when you change colors. This is hardly the first major outage where the AWS status board was a bald-faced lie. This deserves calling out and shaming the responsible parties, nothing less, certainly not defense of blatantly deceptive practices that most companies not named Amazon don't dip into.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#216Problem is that I have to defend our own infrastructure real availability numbers vs cloud's fictional "five nines". It's a loosing game.
Some orgs really do have lousy availability figures (such as my own, the Navy). We have an environment we have access to for hosting webpages for one of the highest leaders in the whole Dept of Navy. This environment was DOWN (not "degrade availability" or "high latencies"), literally off of the Internet entirely, for CONSECUTIVE WEEKS earlier this year. Completely incommunicado as well. It just happened to start wor…
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#217Having an internal network like this that everything on the main AWS network so heavily depends on is just bad design. One does not create a stable high tech spacecraft and then fuels it with coal.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#218Earlier quoted context omitted.
I worked at Amazon. While my boss was on vacation I took over for him in the "Launch readiness" meeting for our team's component of our project. Basically, you go to this meeting with the big decision makers and business people once a week and tell them what your status is on deliverables. You are supposed to sum up your status as "Green/Yellow/Red" and then write (or update last week's document) to explain your stat…
We have something similar at my big corp company. I think the issue is you went from Green to Red in a flip of a switch. A more normal project goes Green...raise a red flags...if red flags aren't resolved in the next week or two, go to yellow...In these meetings everyone collaborates ways to keep your green or get you back to green if you went yellow. In essence - what you were saying is your boss lied the whole time…
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#219So was this in service to something like DynamoDB or some other service?
As in, did some of those extra services that AWS offers for lockin (and that undermines open source projects with embrace and extend) bomb the mainline EC2 service?
Because this kind of smacks of "Microsoft Hidden APIs" that office got to use against other competitors. Does AWS use "special hardware capabilites" to compete against other companies offering roughtly the same service?
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#220My company uses AWS. We had significant degradation for many of their APIs for over six hours, having a substantive impact on our business. The entire time their outage board was solid green. We were in touch with their support people and knew it was bad but were under NDA not to discuss it with anyone. Of course problems and outages are going to happen, but saying they have five nines (99.999) uptime as measured by…
"Our Support Contact Center also relies on the internal AWS network, so the ability to create support cases was impacted from 7:33 AM until 2:25 PM PST. " This to me is really bad. Even as a small company, we keep our support infrastructure separate. For a company of Amazon's size, this is a shitty excuse. If I cannot even reach you as a customer for almost 7 hours, that is just nuts. AWS must do better here. Also, i…
It turned out our call center supplier had something running on AWS, and it took out our entire phone system. After this situation settles, I'm tempted to ask my supplier to see what they're doing to get around this in the future, but I doubt even they knew that AWS was used further downstream.
AWS operates a lot like Amazon.com, the marketplace now--you can try to escape it, but it's near impossible. If you want to ban usage of Amazon's services, you're going to find some service (AWS) or even a Shopify site (FBA warehouse) who uses it.