Details here: http://travaux.ovh.net/?do=details&id=28244 Apparently, the root cause of that issue is a critical software bug in Cisco NCS 2000 transponders.
Their interfaces lost their configuration, and they re-applied configuration, and state came back. This does not equal critical software bug.
> "One of the solutions is to create 2 optical node systems instead of one. 2 systems, that means 2 databases and so in case of loss of configuration, only one system is down. If 50% of the links go through one of the systems, today we would have lost 50% of the capacity but not 100% of links."
This is a crap mitigation. They're still depending on the same hardware and process that led to the first outage, only now there's more of it, so there's more chances to fail.
If they had continuous configuration automation they would have detected when the router's state changed, identified the missing bits, and applied configuration.
"New" routers (as in, since 2011) have APIs and can even run code directly on the router in order to fulfill these requirements. Cisco has multiple white papers, and even provides complete products to manage and certify configuration is applied as desired, even in cloud-agnostic multi-tier networks. Even on old routers, practically all config management solutions out there have plugins to manage Cisco routers.
It's also ridiculous that they had no access to remote hands. This is IT 101.