A maintenance process whose purpose in life is to delete data from ELB backend database (if it were not the case, you'd see "maint process didn't work right") in such a weird way that it would cause such chaos? Why on earth would such a maint process exist in the first place? I can imagine some possibilities here of course but it feels to me there is more to it than what they've chosen to disclose.
Next. So they lose config data but data path for now not impacted. Fine, makes sense. But then backend, with only partial data, attempts to reconfig running LBs, doesn't fail completely (as in it was able to connect and do at least something but not all actions it was supposed to do) thus forcing otherwise good but impaired LBs into a completely bad state. Sunds suspicious to me.
And then the biggest question - why did they choose to attempt to restore entire backend database when only 6.8% of LBs were impacted?
I also have no idea how a CM process can protect against making a mistake - mistakes happen when somebody is at the controls with or without prior coordination.
All in all, their backend systems are so sophisticated and precisely engineered that any unforeseen/unexpected abnormality caused by manual intervention (be it inadvertent run of a maint script or fat fingered traffic re-routing from primary to backup network) inevitably lead to overwhelming reaction of their automation that makes the problem even worse and extremely hard to recover from.
Very tough position to be in - during these outages, they are essentially fighting the skynet that they themselves created and at their scale there is no way around it.
So hats off to those who've been working on this and good luck taming the beast.