Earlier quoted context omitted.
And apparently they had never tried rebooting some of the most important parts of that system. Just when you start to think that someone's really gotten it right you come to learn they're just fumbling around in the dark like everyone else.
To put what @jonhohle said another way, Amazon had probably never brought up the entirety of S3 from Zero to Production-Ready on in a production environment before. I wouldn't necessarily classify this as "fumbling around in the dark." Perhaps they should have tested this in a simulated environment, but (to be fair) on a distributed fault-tolerant system, it probably wasn't a top-priority situation to test.
What makes you think they didn't?