DNS Outage Post Mortem
github.com
DNS Outage Post Mortem
1–10 of 16 posts
Re: DNS Outage Post Mortem
#2This isn't a GitHub only issue but rather one that would affect all quick-to-launch startups (most). What I'm learning from this is that one needs to regularly revisit the infrastructure and how it's glued together with the provisioning system.
If it's not broken, break it.
Re: DNS Outage Post Mortem
#3Re: DNS Outage Post Mortem
#4Who in the right mind would schedule a critical infrastructure upgrade during the day?
Re: DNS Outage Post Mortem
#5Who in the right mind would schedule a critical infrastructure upgrade during the day?
I doubt there is a time when they wouldn't have disrupted a significant part of their userbase. Even if you assume a specific place has the majority of users (San Fransisco, Germany, whatever) developers tend to work odd hours anyway.
Re: DNS Outage Post Mortem
#6Who in the right mind would schedule a critical infrastructure upgrade during the day?
Re: DNS Outage Post Mortem
#7Who in the right mind would schedule a critical infrastructure upgrade during the day?
What is the difference between day and night when your users are worldwide?
This has the great effect of lowering the median travel times and information transmission latencies between the world's population centers, and it means that for at least this geological epoch; we're always going to have daily global peak and off-peak times for human-driven activity.
Re: DNS Outage Post Mortem
#8Who in the right mind would schedule a critical infrastructure upgrade during the day?
What is the difference between day and night when your users are worldwide?
Re: DNS Outage Post Mortem
#9Who in the right mind would schedule a critical infrastructure upgrade during the day?
Re: DNS Outage Post Mortem
#10I would love more detail in the type II error in this validation step, and is worth exploring deeper. What was the verification step? Why did it not detect the issue? What review process was used for the verification step?
While the failed verification step is not the root cause, having good safety checks are the most important part of planning good changes, whether they're DNS reconfigurations, network changes or software deployments.