So an RE failed and flipped to the secondary. They run mlag from the agg layer to their tors. Said mlag had a grey failure. The time to recovery was predominantly fault detection. After detection they upgraded the os on the known good switch. Presumably this bounced, affecting remaining traffic. After upgrade they downed the failed switch, shifting traffic to the known good. Moderately interesting that they run every…
They're probably learning that MLAG should be avoided unless absolutely necessary. But providing common L2 domains across cabinets is probably something they "need." There's much less to go wrong with ECMP at L3. Stateful networking components frighten me.
Re: Network Instability in NYC2 on July 29, 2014
#21Not knowing anything about their infrastructure Id guess that they're using vlan tags for customer isolation. They'd want customer instances spread among racks, based on instance type etc. Going to in house/vxlan/nvgre encap certainly looks better suited, but still has a high bar to entry.