Earlier quoted context omitted.
Sure, but there is always the possibility that then you shut down trading when things _arent_ broken. There are always two error rates. Defining behavior is great for retrospective analysis but would you really feel comfortable putting hard cuts into production based on the answers to those questions? I’m genuinely asking, because IME I wouldn’t be.
That last nine in a trading system uptime has exponentially low value unless you have customers who care quite a lot. Seriously, suppose you have a truly awesome system making $100B per year of revenue. If you unnecessarily shut down 0.1% of the time, that’s only $100M per year lost, and an 0.1% unnecessary shutdown rate seems pretty high.
IME that last 9 is where all the action happens
> unless you have customers who care quite a lot
All customers care about their trades. I’ve worked with these systems. You can’t treat smaller traders as less-than.
> only $100M
How far removed from the problem do you have to be to think one hundred million dollars is not going to effect anyone?