Live data from Hacker News

Network Instability in NYC2 on July 29, 2014

digitalocean.com

1–10 of 22 posts

Re: Network Instability in NYC2 on July 29, 2014

#2
Good post.

This is essentially the risk you face when dealing with new providers.

I guarantee that other providers had to go thought the same set of issues and phases prior to achieving a truly redundant infrastructure.

Would love to hear a follow up on the audit.

Edit: Misread a part of the post. Thought they were doing fail-over on a different network level.

Re: Network Instability in NYC2 on July 29, 2014

#3
post #2

Good post. This is essentially the risk you face when dealing with new providers. I guarantee that other providers had to go thought the same set of issues and phases prior to achieving a truly redundant infrastructure. Would love to hear a follow up on the audit. Edit: Misread a part of the post. Thought they were doing fail-over on a different network level.

Don't schools teach reading comprehension these days? "... the solid state disk in the active routing engine failed on one of our core switches. This triggered a failover to the backup routing engine..."

Re: Network Instability in NYC2 on July 29, 2014

#4
post #2

Good post. This is essentially the risk you face when dealing with new providers. I guarantee that other providers had to go thought the same set of issues and phases prior to achieving a truly redundant infrastructure. Would love to hear a follow up on the audit. Edit: Misread a part of the post. Thought they were doing fail-over on a different network level.

Pretty common for router/switch manufacturers to include just 1 storage device (SSD/CF/whatever). That's one reason why you buy the second router/switch for redundancy.

Re: Network Instability in NYC2 on July 29, 2014

#5
post #2

Good post. This is essentially the risk you face when dealing with new providers. I guarantee that other providers had to go thought the same set of issues and phases prior to achieving a truly redundant infrastructure. Would love to hear a follow up on the audit. Edit: Misread a part of the post. Thought they were doing fail-over on a different network level.

Don't schools teach reading comprehension these days? "... the solid state disk in the active routing engine failed on one of our core switches. This triggered a failover to the backup routing engine..."

FWIW, this response might've been upvoted rather than downvoted if not for the snark.

Re: Network Instability in NYC2 on July 29, 2014

#6
post #4
post #2

Good post. This is essentially the risk you face when dealing with new providers. I guarantee that other providers had to go thought the same set of issues and phases prior to achieving a truly redundant infrastructure. Would love to hear a follow up on the audit. Edit: Misread a part of the post. Thought they were doing fail-over on a different network level.

Pretty common for router/switch manufacturers to include just 1 storage device (SSD/CF/whatever). That's one reason why you buy the second router/switch for redundancy.

Thanks.

Corrected my post. For some reason I thought failure point was outside of the actual switch appliance.

Re: Network Instability in NYC2 on July 29, 2014

#7
So an RE failed and flipped to the secondary. They run mlag from the agg layer to their tors. Said mlag had a grey failure. The time to recovery was predominantly fault detection. After detection they upgraded the os on the known good switch. Presumably this bounced, affecting remaining traffic. After upgrade they downed the failed switch, shifting traffic to the known good.

Moderately interesting that they run everything off a single agg pair. Also that they use mlag instead of routing/mpls/etc for availability.

Key finding is lack of visibility in to the layer 2 availability and performance. Would be interestin to see if they try to ecmp layer 3 or use existing lacp frames for fault detection in the future.

Re: Network Instability in NYC2 on July 29, 2014

#8
post #7

So an RE failed and flipped to the secondary. They run mlag from the agg layer to their tors. Said mlag had a grey failure. The time to recovery was predominantly fault detection. After detection they upgraded the os on the known good switch. Presumably this bounced, affecting remaining traffic. After upgrade they downed the failed switch, shifting traffic to the known good. Moderately interesting that they run every…

They're probably learning that MLAG should be avoided unless absolutely necessary. But providing common L2 domains across cabinets is probably something they "need."

There's much less to go wrong with ECMP at L3. Stateful networking components frighten me.

Re: Network Instability in NYC2 on July 29, 2014

#9
While the outage was upsetting, the response from DO was reassuring. Not only did they jump right on the problem, they maintained communication during the affected period. Then, they refunded me $160 (a month of service).

Every host has an occasional problem. It's how they handle the problem that is important to me. This is a night and day contrast compared to the service I received with other hosts with which I've dealt.

Re: Network Instability in NYC2 on July 29, 2014

#10
I find it alarming that they don't have a competent network engineer of their own on staff. Note the following:

"We are working very closely with our networking partner to understand the nature of the failure, assess the chances of a repeat event, and to begin planning architectural changes for the future."

"Our initial focus was on verifying the configuration so we initiated a line-by-line configuration review by engineers at our network partner"

Yikes. You get what you pay for.

Post reply on HN