The fact that so many popular sites/services are experiencing issues due to a single AZ failure makes me think that there is a serious shortage of good cloud architects/engineers in the industry. It would be one thing if this was a Regional failure, but a single AZ failure should not have any noticeable effect.
Architect here. We had an outage and we have a very complete architecture. The issue is, the services were still reachable via internal health checks. So instead of taking the effected servers out of service they stayed in. We had to resolve it by manually shutting down all the servers in the affected AZ. Which is normally not needed. There are of course a lot of companies that aren't architected with multi-AZ at all…
P.S. AWS has said they have resolved the issue for almost 2 hours now and we are still having issues with us-east-2a.