Root cause analysis: significantly elevated error rates on 2019‑07‑10
1–10 of 114 posts
Re: Root cause analysis: significantly elevated error rates on 2019‑07‑10
#2Re: Root cause analysis: significantly elevated error rates on 2019‑07‑10
#3Re: Root cause analysis: significantly elevated error rates on 2019‑07‑10
#4This reads like the marketing/PR teams wrote much of it. Compare to the Cloudflare post-mortem from today: https://blog.cloudflare.com/details-of-the-cloudflare-outage...
Re: Root cause analysis: significantly elevated error rates on 2019‑07‑10
#5Re: Root cause analysis: significantly elevated error rates on 2019‑07‑10
#6This reads like the marketing/PR teams wrote much of it. Compare to the Cloudflare post-mortem from today: https://blog.cloudflare.com/details-of-the-cloudflare-outage...
Also, does your company's engineering decisions change based on other companies' post-mortems?
Re: Root cause analysis: significantly elevated error rates on 2019‑07‑10
#7Damn what a mess. Sounds like y'all are rolling out way to many changes too quickly with little to no time for integration testing.
It's a somewhat amateur move to assume you can just arbitrarily rollback without consequence, without testing etc.
One solution I don't see mentioned, don't upgrade to minor versions ever. And create a dependency matrix so if you do rollback, you rollback all the other things that depend on the thing you're rolling back as well.
Re: Root cause analysis: significantly elevated error rates on 2019‑07‑10
#8How did they not catch this? It's super surprising to me that they wouldn't have monitors for this.
Re: Root cause analysis: significantly elevated error rates on 2019‑07‑10
#9This reads like the marketing/PR teams wrote much of it. Compare to the Cloudflare post-mortem from today: https://blog.cloudflare.com/details-of-the-cloudflare-outage...
Re: Root cause analysis: significantly elevated error rates on 2019‑07‑10
#10This reads like the marketing/PR teams wrote much of it. Compare to the Cloudflare post-mortem from today: https://blog.cloudflare.com/details-of-the-cloudflare-outage...
I'm Stripe's CTO and wrote a good deal of the RCA (with the help of others, including a lot of the engineers who responded to the incident). If you've any specific feedback on how to make this more useful, I'd love to hear it.
If anyone on HN knows anyone who has the sort of interesting life story where they both know what can cause a cluster election to fail and like writing about that sort of thing, we would eagerly like to make their acquaintance.