Live data from Hacker News

GoDaddy outage caused by corrupted router tables

godaddy.com

71–80 of 97 posts

Re: GoDaddy outage caused by corrupted router tables

#72
post #20

I find this extremely suspicious (I.E. knowing routers, I call bullshit). The change to the Verisign anycast DNS service which I noted yesterday in another thread... brought godaddy.com back up, yet did not result in bringing other DNS services back up. Someone is lying here in my opinion. I hope I'm proven wrong because this is a terrible excuse for the company to make. EDIT: And as someone else pointed out... their…

What we're being told internally matches the public response. Reading between the lines, it sounds like there may have been human error involved, but that's just speculation.

Our interim CEO confirmed that affected users are receiving a full month refund.

Re: GoDaddy outage caused by corrupted router tables

#73
This statement makes a lot of sense. I found it a bit suspicious that Anonymous Own3r twitted: "When i do some DDOS attack i like to let it down by many days, the attack for unlimited time, it can last one hour or one month" Which sounds like he actually has no control over what is happening and makes a statement that is impossible to disprove.

Re: GoDaddy outage caused by corrupted router tables

#75
post #18

Earlier quoted context omitted.

Routers at this level aren't just scale-ups of your home wifi/nat box. They aren't even scale ups of the simple IP routers for a basic IT data-closet that manages subnets and whatnot (already much more complex by dealing with vlans and subnets and dmz and vpn issues). At the level of big networking company they are a truly complex beast. Just at the IP level they have to deal with (at the edges and across substantial…

...and if you don't have up to date backups of your router tables, it will take a long time to recover from an "oops". Doing the wrong thing to router table(s) is the network equivalent of "sudo rm -rf /".

Any network engineer i've ever worked with that's been worth his salt has had a system in place to backup his router configs before he does any update. The last place I worked at, we had a program setup to automatically tftp download the router configs at night and check them in to source control. We could then run diffs to see what changed if there was some kind of outage.

Re: GoDaddy outage caused by corrupted router tables

#76
"yet did not result in bringing other DNS services backup"

Can you be more specific?

Which other domain names did you try?

Also, I believe some parts of the world were unaffected by the outage.

I would guess a large majority of GoDaddy customers would not even know this outage occurred. They are "casual" domain name registrants and in some cases "casual" website operators. They registered some names and then never did anything with them. Or they operate a website but it's very low traffic and they rarely think about it. That is only a guess.

Re: GoDaddy outage caused by corrupted router tables

#77
post #72
post #20

I find this extremely suspicious (I.E. knowing routers, I call bullshit). The change to the Verisign anycast DNS service which I noted yesterday in another thread... brought godaddy.com back up, yet did not result in bringing other DNS services back up. Someone is lying here in my opinion. I hope I'm proven wrong because this is a terrible excuse for the company to make. EDIT: And as someone else pointed out... their…

What we're being told internally matches the public response. Reading between the lines, it sounds like there may have been human error involved, but that's just speculation. Our interim CEO confirmed that affected users are receiving a full month refund.

And if it was human error, that's fine! Stuff happens and I can certainly say that I've done my share of human error.

I want to mention a couple things though. The first is that the blame is being placed (at least reading slightly into the PR release) on a 'technology' failure. That is fairly distinct from human error.

The other thing, is that if it was human error, how did the chain occur without a second pair of eyes or similar, such that the outage lasted more than a bit of time?

Third, why was the DNS changed to Verisign? That is still the biggest outstanding question I think in terms of their claimed outage reports. I should also mention that I do have skin in this game as plenty of customers were running at least SOMETHING through Godaddy and were yelling in this direction for stuff breaking...

Re: GoDaddy outage caused by corrupted router tables

#78

Earlier quoted context omitted.

Ex-GoDaddy employee here: Everyone there is on pins and needles at this point. Since Silverlake's investment in the company, many hatchets have dropped on jobs, and it's really the only decent tech firm in Phoenix to work at. My guess is that there is some hiney covering going on with this explanation, and the interim CEO has little cause to care too much about responsibility, since he'll likely be out before year's…

can you clarify what you mean by: > it's really the only decent tech firm in Phoenix to work at. and then: > Horrible, horrible, backward-ass culture. what part of the company is good if not the culture?

There are many tech (and non-tech) companies in the Phoenix area. Your choice, as always, is to work at a place where the importance of the technology and the talent of your engineers and designers is recognized, or work in PetSmart's IT/Web department (where you're constantly fighting budget and recognition battles).

So should you opt to go into tech, GoDaddy is the largest and best-paying, and has fairly decent benefits. Unless you were a Senior level professional elsewhere, it would likely be in your best interest to work at GD.

That said, it's also a company run by bean counters and marketing. Most (read: all) important decisions regarding what choices the company makes goes through a ringer that includes assessing how much direct money comes from an innovation or change. If test a yields or saves $1 and test b yields or saves $2, test b wins, no matter how poor a choice it is in terms of user experience, customer care or any other metrics that relate to long-term customer retention.

It's got tons of middle management, which in itself isn't a bad thing, except that everyone's fighting to own the creation a product, but no one wants to be accountable for it, should it not go well.

Essentially this creates an environment of fear against innovation, accountability and iteration.

Re: GoDaddy outage caused by corrupted router tables

#80

Earlier quoted context omitted.

Ex-GoDaddy employee here: Everyone there is on pins and needles at this point. Since Silverlake's investment in the company, many hatchets have dropped on jobs, and it's really the only decent tech firm in Phoenix to work at. My guess is that there is some hiney covering going on with this explanation, and the interim CEO has little cause to care too much about responsibility, since he'll likely be out before year's…

> "It won't fail because of me." Leave now. It will be a black mark on your resume if you stay. Seriously.

> It will be a black mark on your resume if you stay.

It's funny you say this. This is exactly the reason why I left, and exactly what I told HR when I left. Unsurprisingly, they had told me that others who had recently left gave similar sentiment.

Post reply on HN