Live data from Hacker News

WAN router IP address change blamed for global Microsoft 365 outage

theregister.com

51–60 of 68 posts

Re: WAN router IP address change blamed for global Microsoft 365 outage

#52
post #14

The curse of network engineering. You’re invisible and insignificant when everything is running well, and public enemy number one if you make a mistake!

this is the general case with all critical systems. Everything from networking to sewers (... not actually that different now that I mention it) to pandemic planning. No one gets credit for the pandemic prevented because the BSL regulations did their job.

Re: WAN router IP address change blamed for global Microsoft 365 outage

#53
post #45

Earlier quoted context omitted.

My point is that you could apply the same principle internally; have two backbones managed by separate teams instead of one.

That still doesn't make sense though. In the context of a WAN, a backbone is an external network. It routes between your POPs. At any rate, the margin of error and complexity in having two separate backbones networks managed by two separate teams would likely result in more network issues not less. The whole point in having an AS is having a coherent routing policy.

The parent never said multiple networks was easier to implement.

In fact it could easily 2x the cost for the same level of quality, which is why it's almost unheard of for cloud.

Re: WAN router IP address change blamed for global Microsoft 365 outage

#54

Earlier quoted context omitted.

Redundant vendors in the GP’s context referred to using multiple router vendors, eg Cisco and Juniper. Using multiple connectivity vendors doesn’t guarantee path diversity. Demanding fibre maps and ensuring that your connectivity has separate points of entry into the building, doesn’t cross outside the building, and validating with your DC provider that your cross connects aren’t crossing either, guaranteed path dive…

The GP was clearly talking about whole networks, not just the hardware vendors, if I read that different than the GP intended I'll wait for their correction. One of the problems that I've seen in practice that with the degree of virtualization at play that it has at the same time become much more easy to in principle be guaranteed 100% independence and in practice it has become much harder to verify that this is the…

Sounds like a great place for a specialized insurance company to be the middle man

Re: WAN router IP address change blamed for global Microsoft 365 outage

#55

Earlier quoted context omitted.

That still doesn't make sense though. In the context of a WAN, a backbone is an external network. It routes between your POPs. At any rate, the margin of error and complexity in having two separate backbones networks managed by two separate teams would likely result in more network issues not less. The whole point in having an AS is having a coherent routing policy.

The parent never said multiple networks was easier to implement. In fact it could easily 2x the cost for the same level of quality, which is why it's almost unheard of for cloud.

The parent was stating that two networks would be better but its none done because of costs. And that's complete nonsense.

The fact that it's more difficult and complex to have two separate teams manage two separate networks means it's more prone to error and misconfiguration. The reason it's not done has nothing do with financial costs but rather because it makes no sense, for the very fact I just mentioned.

Re: WAN router IP address change blamed for global Microsoft 365 outage

#56
post #45

Earlier quoted context omitted.

My point is that you could apply the same principle internally; have two backbones managed by separate teams instead of one.

That still doesn't make sense though. In the context of a WAN, a backbone is an external network. It routes between your POPs. At any rate, the margin of error and complexity in having two separate backbones networks managed by two separate teams would likely result in more network issues not less. The whole point in having an AS is having a coherent routing policy.

You're talking like multihoming doesn't work. Sure there are cases where bugs or bad configs can propagate across ASes but most of the time you can survive if one provider goes down.

Re: WAN router IP address change blamed for global Microsoft 365 outage

#57

Earlier quoted context omitted.

The parent never said multiple networks was easier to implement. In fact it could easily 2x the cost for the same level of quality, which is why it's almost unheard of for cloud.

The parent was stating that two networks would be better but its none done because of costs. And that's complete nonsense. The fact that it's more difficult and complex to have two separate teams manage two separate networks means it's more prone to error and misconfiguration. The reason it's not done has nothing do with financial costs but rather because it makes no sense, for the very fact I just mentioned.

Two end to end networks would be more reliable.

Like two independent internets spanning from your server to my laptop.

Two completely isolated end to end transports.

That's what OP meant when they said you could make it more reliable by having a redundant network. It's just prohibitively expensive.

Then if one internet goes down in any way I talk to you over the other. That's a fairly straightforward fallback algorithm to implement.

Re: WAN router IP address change blamed for global Microsoft 365 outage

#58
post #16

Earlier quoted context omitted.

This may be true of enterprise network engineers but I’ve worked across a lot of very large networks (telco, not cloud) and we never ever trust the vendor. The kind of bugs that I’ve read about in errata notes over the years is wild and truly unpredictable.

It wouldn't be the first time that your redundant vendors end up sharing a conduit for a bunch of fiber somewhere. Guess where that backhoe will start digging?

I have to trust the dark fibre map provided, but I know exactly which way it ran, manhole to manhole. I had three cores, they shared the first 20 metres to the manhole, it's unlikely there would be a backhoe digging underneath the police van and pile of scaffolding that was parked in the shared conduit.

After that it went on different paths to three different buildings, which from each of those was then routed independently.

We take physical resilience seriously, as it isn't network engineers that do that part of the infrastructure. Enterprise network engineers then throw it all away by stacking their switches into a single point of logical failure.

(Still had a non-IP backup, but sometimes that breaks too -- just in different ways than the IP)

Re: WAN router IP address change blamed for global Microsoft 365 outage

#59

Earlier quoted context omitted.

Redundant vendors in the GP’s context referred to using multiple router vendors, eg Cisco and Juniper. Using multiple connectivity vendors doesn’t guarantee path diversity. Demanding fibre maps and ensuring that your connectivity has separate points of entry into the building, doesn’t cross outside the building, and validating with your DC provider that your cross connects aren’t crossing either, guaranteed path dive…

The GP was clearly talking about whole networks, not just the hardware vendors, if I read that different than the GP intended I'll wait for their correction. One of the problems that I've seen in practice that with the degree of virtualization at play that it has at the same time become much more easy to in principle be guaranteed 100% independence and in practice it has become much harder to verify that this is the…

In London I can literally follow the map from manhole to manhole, exchange to exchange. It's dark fibre so I can flash a light down it and a colleague can see it emerge at the other end. Now it's possible they don't follow the map and still make it to the other end, but it's pretty unlikely.

Sometimes of course you have to make judgement calls. From one location near Slough I have a BT EAD2 back to my building a few miles away. I know the route into my building, I can see the cables with my own eyes going in different directions. BT tell me which exchanges those cables goto, and provide me with a map into the field at a 1000:1 scale showing the cables coming in down a shared path. Sure it's possible BT are lying, but it's unlikely. Only use that location sporadically, and when I do it's a managed location, so I can accept the risk of a digger on the ground.

Another location in Norfolk, two BTNet lines, going to two different exchanges. They meet at the edge of the farm and go up the same trunk. That's fine, I can physically control the single point of failure there too, although if peering between BT and my network fails then I'm screwed, but I have a separate pinnacom circuit in a crunch.

Now obviously some failure become far harder to mitigate. A failure of the Thames Barrier would cause a hell of a lot of problems in Docklands, I'm not sure if any circuits in/out of places like telehouse, sovhouse, etc will remain. Cross that bridge etc. Whether my electricity provider will remain with a loss of the internet is another matter, so then it comes down to how much oil there in in the generators, and the generators of any repeaters on the routes of my network.

However the much easier to avoid is the problem of some shitty stacked switch the salesman says will always work.

Re: WAN router IP address change blamed for global Microsoft 365 outage

#60

Earlier quoted context omitted.

I remember the first time I got access to an employers production Cisco router. It’s pretty scary how easy it is to majorly fuck something up. There isn’t a concept of a transaction or a rollback. You just enter a command, press enter and it’s live. To counter this we’d write all the commands we planned on executing and peer review it. Nothing was to be done “on the fly” (at least in theory) In short, coming from a d…

> There isn’t a concept of a transaction or a rollback. Yeah, Cisco gear is bonkers. Mikrotik has "Safe Mode", which undoes all commands since you entered "Safe Mode" if the connection that created the shell gets interrupted. It has saved my bacon on several occasions, but there are several obvious situations in which you can get yourself locked out. Juniper gear has "commit confirmed $NUMBER_OF_MINUTES", which will…

> I do have no idea how Juniper's rollback works when multiple users are doing simultaneous config editing... maybe don't do that?

You get a warning

    Users currently editing the configuration:
      bob termainal p0...."
But the failure here is actually sshing to a network switch in the first place.

Some cisco kit has restconf which is better for automation, but it's buggy.

Post reply on HN