> First of all tons of more outages than ever, remember that React useEffect fkup [0]? Complete insane that this would happen at an infra company that runs a third or so of the web. Bugs happen all the time. They roughly increase with scale, not decrease. There’s an argument to be made about better testing, but this specific bug seems like a perfect one to slip through: multiple services, hard to spot at code review,…
Bugs happen all the time yes, outages don't. If outages are increasing with scale you get 1 or maybe 2 free passes. After that you either have in-ept Engineering or just in-ept leadership. I used to be all in on CF a while ago, now I am moving off them almost entirely. Same issue with GitHub, I can understand if you can't build for the scale when you couldn't predict it but if after over 12-18 months things don't see…
AWS has had more than 1 outage a year - is it not worth investing in? Are they not serious?
Outages are just a specific kind of bug, often surfaced by the interactions of several discrete bugs.
Saying “you’re not serious if you have more than 1 bug a year” is silly.