Earlier quoted context omitted.
I've been naively setting up our distributed databases in separate AZs for a couple years now, paying, sometimes, thousands of dollars per month in data replication bandwidth egress fees. As far as I can remember I've never never seen an AZ go down, and the only region that has gone down has been us-east-1.
Is that separate AZs within the same region, or AZs across regions? I didn't think there were any bandwidth fees between AZs in the same region.
Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
101–110 of 410 posts
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#102Obviously one hopes these things don’t happen, but that’s an impressive and transparent write up that came out quickly.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#103My favorite sentence: "Our networking clients have well tested request back-off behaviors that are designed to allow our systems to recover from these sorts of congestion events, but, a latent issue prevented these clients from adequately backing off during this event."
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#104Earlier quoted context omitted.
I mean, not to defend them too strongly, but literally half of this post mortem is addressing the failure of the Service Dashboard. You can take it on bad faith, but they own up to the dashboard being completely useless during the incident.
Multiple AWS employees have acknowledged it takes VP approval to change the status color of the dashboard. That is absurd and it tells you everything you need to know. The status page isn't about accurate information, it's about plausible deniability and keeping AWS out of the news cycle.
How'd that work out for them?
https://duckduckgo.com/?q=AWS+outage+news+coverage&t=h_&ia=w...
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#105My company uses AWS. We had significant degradation for many of their APIs for over six hours, having a substantive impact on our business. The entire time their outage board was solid green. We were in touch with their support people and knew it was bad but were under NDA not to discuss it with anyone. Of course problems and outages are going to happen, but saying they have five nines (99.999) uptime as measured by…
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#106My company uses AWS. We had significant degradation for many of their APIs for over six hours, having a substantive impact on our business. The entire time their outage board was solid green. We were in touch with their support people and knew it was bad but were under NDA not to discuss it with anyone. Of course problems and outages are going to happen, but saying they have five nines (99.999) uptime as measured by…
I mean, not to defend them too strongly, but literally half of this post mortem is addressing the failure of the Service Dashboard. You can take it on bad faith, but they own up to the dashboard being completely useless during the incident.
Let's not act like this is the first time this has happened. It's bad faith that they do not change when their promise is they hire the best to handle infrastructure so you don't have to. It's clearly not the case. Between this and billing I we can easily lay blame and acknowledge lies.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#107Earlier quoted context omitted.
Take DynamoDB as an example. The AWS managed service takes care of replicating everything to multiple AZs for you, that's great! You're very unlikely to lose your data. But, the DynamoDB team is running a mostly-regional service. If they push bad code or fall over it's likely going to be a regional issue. Probably only the storage nodes are truly zonal. If you wanted to deploy something similar, like Cassandra across…
> But, the DynamoDB team is running a mostly-regional service. this is both more and less true than you might think. for most regional endpoints teams leverage load balancers that are scoped zonally, such that ip0 will point at instances in zone a, ip1 will point at instances in zone b, and so on. Similarly, teams who operate "regional" endpoints will generally deploy "zonal" environments, such that in the event of a…
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#108Earlier quoted context omitted.
I mean, not to defend them too strongly, but literally half of this post mortem is addressing the failure of the Service Dashboard. You can take it on bad faith, but they own up to the dashboard being completely useless during the incident.
People of HN has been extremely unprofessional with regards to AWS's downtime. Some kind of a massive zeitgeist against Amazon, like a giant hive mind that spews hate. Why are we doing this folks? What's making you so angry and contemptful? Literally try searching the history of downtimes and it was always professional and respectful. Yesterday, my comment was fricking flagged for asking people to be nice to which pe…
Because Amazon kills industries. Takes job. They do this because they promise they hire the best people that can do this better than you and for cheaper. And it's rarely true. And then they lie about it when things hit the fan. If you're going to be the best you need to act like the best, and execute like the best. Not build a walled garden that people cant see into, and hard to leave.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#109Having an internal network like this that everything on the main AWS network so heavily depends on is just bad design. One does not create a stable high tech spacecraft and then fuels it with coal.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#110My company uses AWS. We had significant degradation for many of their APIs for over six hours, having a substantive impact on our business. The entire time their outage board was solid green. We were in touch with their support people and knew it was bad but were under NDA not to discuss it with anyone. Of course problems and outages are going to happen, but saying they have five nines (99.999) uptime as measured by…
I worked at Amazon. While my boss was on vacation I took over for him in the "Launch readiness" meeting for our team's component of our project. Basically, you go to this meeting with the big decision makers and business people once a week and tell them what your status is on deliverables. You are supposed to sum up your status as "Green/Yellow/Red" and then write (or update last week's document) to explain your stat…
1) If you keep status green for 5 years, while not delivering anything, the reality is the folks at the very top (who can come and go) just look at these colors and don't really get into the project UNLESS you say you are red :)
2) Within 1-2 years there is always going to be some excuse for WHY you are late (people changes, scope tweaks, new things to worry about, covid etc)
3) Finally you are 3 years late, but you are launching. Well, the launch overshadows the lateness. Ie, you were green, then you launched, that's all the VP really sees sometime.