Recently one of my production node landed on a bad physical host, DO reboot my node twice within 24 hours, I tweeted about it and DO support open a ticket for me. Since the physical node have problem, I plan to take snapshot then spin up a new node using the image. However, as the physical node have problem, the snapshot take more than 7 hours and still failed to snapshot (No way to cancel snapshot task). DO support…
Plan for your server to go down at any time, and, if you have actual $$$ in play, also plan for the zone to go down.
Amazon encourages this thinking by having very low-latency links between AZ (Availability Zones) - and they also recommend deploying in two geographical regions should a regional disaster occur.