Live data from Hacker News

Level 3 Global Outage

puck.nether.net

211–220 of 393 posts

Re: Level 3 Global Outage

#212
post #124

Massive reconvergence event in their network, causing edge router bgp sessions to bounce (due to cpu). Right now all their big peers are shutting down sessions with them to give level3s network the ability to reconverge. Prefixes announced to 3356 are frozen on their route reflectors and not getting withdrawn. Edit: if you are a Level3 customer shut your sessions down to them.

History doesn't repeat, but it rhymes .... There was a huge AT&T outage in 1990 that cut off most US long distance telephony (which was, at the time, mostly "everything not within the same area code"). It was a bug. It wasn't a reconvergence event, but it was a distant cousin: Something would cause a crash; exchanges would offload that something to other exchanges, causing them to crash -- but with enough time for th…

There was something similar a few years ago on a large US mobile network. You could watch the ‘storm’ rolling across the map. Fascinating stuff

Re: Level 3 Global Outage

#213
post #202
post #160

Earlier quoted context omitted.

> continue where they left off The games are timed and this pause gives a lot of thinking time. If they're allowed to talk with others during the pause, then also consulting time. > why don't they start over That would be unfair to the player who was ahead. That said, both players might still be fine with a clean rematch, because being the undisputed winner feels better. I wonder if they were asked (anonymously to pr…

Seems like one of those cases where solving a “little” issue would actually require rearchitecting the entire system. Namely, in this case, it seems like the “right thing” is for games to not derive their ELO contributions from pure win/loss/draw scorings at all, but rather for games to be converted into ELO contributions by how far ahead one player was over the other at the point when both players stopped playing fo…

> but rather for games to be converted into ELO contributions by how far ahead one player was over the other at the point when both players stopped playing for whatever reason

Except for the obvious positions that no one serious would even play, there is no agreed-upon way of calculating who has an advantage in chess like that. One man's terrible mobility and probable blunder is another's brilliant stratagem.

Re: Level 3 Global Outage

#215
CenturyLink/Level3 on Twitter: "We are able to confirm that all services impacted by today’s IP outage have been restored. We understand how important these services are to our customers, and we sincerely apologize for the impact this outage caused."

https://twitter.com/CenturyLink/status/1300089110858797063

Re: Level 3 Global Outage

#216
This had me really confused until I saw it was a global outage. I have been getting delayed iOS push notifications (from prowl) now for the last few hours, from a device I was fairly sure I had disconnected 3 hours ago (a pump)

Got questioning if I really disconnected it before I left.

I'm wondering if we're at the point where internet outages should have some kind of (emergency) notification/sms sent to _everyone_.

Re: Level 3 Global Outage

#217

Earlier quoted context omitted.

I had the same issue on my fiber connection (Altibox/BKK), however, no problems on my mobile using 4G (Dipper/Telenor)

I couldn't reach HN on neither Altibox or 4g/telenor.

Both altibox and telia 4g was down for me as well.

Re: Level 3 Global Outage

#218
post #124

Massive reconvergence event in their network, causing edge router bgp sessions to bounce (due to cpu). Right now all their big peers are shutting down sessions with them to give level3s network the ability to reconverge. Prefixes announced to 3356 are frozen on their route reflectors and not getting withdrawn. Edit: if you are a Level3 customer shut your sessions down to them.

What is a reconvergence event? Is that what's described in your last sentence?

https://en.wikipedia.org/wiki/Convergence_(routing)

IP network routing is distributed systems within distributed systems. For whatever reason the distributed system that is the CenturyLink network isn't "converging", or we could it becoming consistent, or settling, in a timely manner.

Re: Level 3 Global Outage

#219
post #115

Earlier quoted context omitted.

Sadly, in my experience, ipv4 is generally more reliable than ipv6 still. Set up two hosts, host A and host B in two different data centers. Make them send HTTP requests to each other over ipv4 and over ipv6. You'll see that latency spikes, packet loss is more frequent over ipv6.

Why is that?

We’ve observed this in end-user devices, especially on some ISPs.

It makes sense if the overall adoption and resource allocation are comparatively smaller, making individual or small-group coincident spikes more impactful against the amortized whole.

It’s a lot like a market with low volume/liquidity. Someone wanders in with a big transaction and blows everything up.

Re: Level 3 Global Outage

#220
post #56
post #55

Internet infrastructure is broken. Why do a few companies control the backbone of the internet? Shouldn’t there be a fallback or disaster recovery plan if one or more of these companies become unavailable?

Why doesn't stuff just route around this automatically, if one provider has problems?

Things mostly routed around the problem. Issues arose because a) some people are single-homed to Level3/CenturyLink b) apparently Level3/CenturyLink continued announcing unreachable prefixes, which breaks the Internet BGP trust model.
Post reply on HN