Live data from Hacker News

GoDaddy outage caused by corrupted router tables

godaddy.com

31–40 of 97 posts

Re: GoDaddy outage caused by corrupted router tables

#31
post #10

Based on a long history of working in datacenters I'd bet someone misconfigured something and later claimed it was "corrupted" to save their ass - happens all the time. It's just so simple to make very confusing and damaging mistakes in a complicated network. I wouldn't be surprised to hear that GoDaddy's corporate culture wouldn't respond well to someone admitting to a mistake this damaging.

Ex-GoDaddy employee here:

Everyone there is on pins and needles at this point. Since Silverlake's investment in the company, many hatchets have dropped on jobs, and it's really the only decent tech firm in Phoenix to work at.

My guess is that there is some hiney covering going on with this explanation, and the interim CEO has little cause to care too much about responsibility, since he'll likely be out before year's end anyway.

I really feel for the folks who work there. Many, many talented people who don't have an inch to make a mistake. When I was last there, they had just released an internal communication about the new company motto: "It won't fail because of me." Horrible, horrible, backward-ass culture.

Re: GoDaddy outage caused by corrupted router tables

#32

How can they claim 99.999% uptime, when they just had several hours of service outage? I'm not sure how long they've been providing DNS hosting, but by the most generous assumption this would be the entire 15 years of their existence. 99.999% allows them about 1.3 hours of outage in 15 years.

That's not how SLAs work.

Basically they're guaranteeing 99.999% uptime across some interval and you get some kind of compensation if/when they fail to meet it.

Re: GoDaddy outage caused by corrupted router tables

#34
post #29

Earlier quoted context omitted.

Unless you have intimate knowledge of their network topology, and know the specifics of where those pinged IP's live in that topology, and what routes were used to provide DNS results, you can't say that it wasn't a routing issue. "Routing" is a rather generic term when it comes to large networks, and everything from border routers, firewalls, load balancers, and switches actually perform routing. Especially (as I've…

None of that is necessarily incorrect... but per their news release 'corrupted router data tables' (their words) were the issue. I can't read too much into that, but that still doesn't change the fact that DNS wasn't resolving for a while after they made their Verisign change (for clients), yet their website was resolved with this change. You are correct that I don't know the details of their internal network and I n…

I can imagine that they'd understandably work to get their own site/etc up and running first as the priority, as a manual "hack". After all, it's the main page everyone would be going to for information on what's going on.

After that, coming up with an automated process for migrating what must be a shit-ton of zone information to another system must have taken some time. I have no idea what their specific solution was, but I'm fairly confident in the fact that it wasn't just a matter of copying over a few zone files. They'd probably have to do SOME sort of ETL (extraction / translation / load) process that would take some time to develop, test, never mind run.

And I can't remember the last time I gave technical information to a PR person who actually got it 100% technically correct. ;)

My intention wasn't to shit on your point or anything, or in any way defend Go-Daddy and their screwup, I'm just thinking it's a bit unrealistic to try and infer detailed information from a PR release.

In the end, it was technical, they screwed up, and I doubt they'd ever release a proper, detailed post-mortem of what happened.

Re: GoDaddy outage caused by corrupted router tables

#35

How can they claim 99.999% uptime, when they just had several hours of service outage? I'm not sure how long they've been providing DNS hosting, but by the most generous assumption this would be the entire 15 years of their existence. 99.999% allows them about 1.3 hours of outage in 15 years.

That's not how SLAs work. Basically they're guaranteeing 99.999% uptime across some interval and you get some kind of compensation if/when they fail to meet it.

I don't believe they have a 99.999% SLA. That was their historical up-time.

I just misread their statement. They probably meant the up-time before this incident.

"Throughout our history, we have provided 99.999% uptime in our DNS infrastructure. This is the level our customers expect from us and the level we expect of ourselves. We have let our customers down and we know it."

Re: GoDaddy outage caused by corrupted router tables

#36
post #20

I find this extremely suspicious (I.E. knowing routers, I call bullshit). The change to the Verisign anycast DNS service which I noted yesterday in another thread... brought godaddy.com back up, yet did not result in bringing other DNS services back up. Someone is lying here in my opinion. I hope I'm proven wrong because this is a terrible excuse for the company to make. EDIT: And as someone else pointed out... their…

Unless you have intimate knowledge of their network topology, and know the specifics of where those pinged IP's live in that topology, and what routes were used to provide DNS results, you can't say that it wasn't a routing issue. "Routing" is a rather generic term when it comes to large networks, and everything from border routers, firewalls, load balancers, and switches actually perform routing. Especially (as I've…

While it could indeed be a routing issue, who's to say that it wasn't caused intentionally by the guy in the tweets? It would be in GoDaddy's interests to cover that up and fix whatever exploit he used to get in, instead of admitting a security breach.

Re: GoDaddy outage caused by corrupted router tables

#37
post #29

Earlier quoted context omitted.

None of that is necessarily incorrect... but per their news release 'corrupted router data tables' (their words) were the issue. I can't read too much into that, but that still doesn't change the fact that DNS wasn't resolving for a while after they made their Verisign change (for clients), yet their website was resolved with this change. You are correct that I don't know the details of their internal network and I n…

I can imagine that they'd understandably work to get their own site/etc up and running first as the priority, as a manual "hack". After all, it's the main page everyone would be going to for information on what's going on. After that, coming up with an automated process for migrating what must be a shit-ton of zone information to another system must have taken some time. I have no idea what their specific solution wa…

Heh, yeah. It is a bit difficult to interpret PR speak (and I have had to correct our guy before).

I think perhaps the takeaway from here is to not trust what is being said, go with your gut... and move any services off GoDaddy ;). Would be nice if like Google or Amazon they would release a real post-mortem post. Even if it's an internal 'uh-oh' I trust companies that are willing to admit to mistakes.

Re: GoDaddy outage caused by corrupted router tables

#38
I used to think GoDaddy went down through no active fault of their own. Now I realize they're just a shitty company with a shitty product. Good riddance.

Our company's email is hosted on GoDaddy, and it has the most downtime of any email service I use (which are Google Apps/Gmail, Yahoo for Business and a private IMAP server). Frequently it will just be dead in the middle of the day, which is why this particular outage wasn't so surprising to us.

Re: GoDaddy outage caused by corrupted router tables

#39
post #10

Based on a long history of working in datacenters I'd bet someone misconfigured something and later claimed it was "corrupted" to save their ass - happens all the time. It's just so simple to make very confusing and damaging mistakes in a complicated network. I wouldn't be surprised to hear that GoDaddy's corporate culture wouldn't respond well to someone admitting to a mistake this damaging.

Ex-GoDaddy employee here: Everyone there is on pins and needles at this point. Since Silverlake's investment in the company, many hatchets have dropped on jobs, and it's really the only decent tech firm in Phoenix to work at. My guess is that there is some hiney covering going on with this explanation, and the interim CEO has little cause to care too much about responsibility, since he'll likely be out before year's…

can you clarify what you mean by:

> it's really the only decent tech firm in Phoenix to work at.

and then:

> Horrible, horrible, backward-ass culture.

what part of the company is good if not the culture?

Post reply on HN