Are there any case studies where microservices went well? From an end user perspective, Netflix runs in “constantly degraded” mod. From an engineering perspective, they track “number of successful stream starts”, instead of percentage of the time 100% of their services are working. That’s a huge red flag. As a researcher, the monitoring and fault-propagation / modeling work they’ve done to get it to stay up at all is…
That doesn't seem true. I would imagine that at Netflix scale, you probably have request tracing libraries that can give you a graph of service dependencies. Whether it's worthwhile to consume that, or easier to just let Chaos Monkey run rampant is another question.
Also, I very rarely have issues with Netflix. Typically when I do I can just exit the stream and restart it. Anecdata, but I could count on one hand the number of times I've had a title just not play, or Netflix be down entirely.