> (...) [W]e have not completely restarted the index subsystem or the placement subsystem in our larger regions for many years. S3 has experienced massive growth over the last several years and the process of restarting these services and running the necessary safety checks to validate the integrity of the metadata took longer than expected. All those tweets saying "turn it off and back on again"? "We accidentally tu…
The system is a collection of shards. If you replicate it to create a second shard, then you'll just have a large a system, which is still a single point of failure.
The index, by necessity, has to be able to answer the question 'this object exists' or 'this object doesn't exit' - so it needs to have consensus.