Live data from Hacker News

Level 3 Global Outage

puck.nether.net

51–60 of 393 posts

Re: Level 3 Global Outage

#53

Shameless plug: I spent too much time losing precious time when github/npm/cloudflare are going down, until I figure out it was them. So currently working on a project[1] to monitor all the 3rd party stack you use for your services. Hit me up if you want, access I'll give free access for a year+ to some folks to get feedbacks. [1] https://monitory.io

Another typo:

> Know when services you depend on goes down

"Services go down", not "goes".

Re: Level 3 Global Outage

#54
post #7
post #5

Earlier quoted context omitted.

Is there any one place that would be a good first place to go to check on outages like this? It would be really cool and useful to have an "public Internet health monitoring center"... this could be a foundation that gets some financing from industry that maintains a global internet health monitoring infrastructure and a central site at which all the major players announce outages. It would be pretty cheap and have a…

Until that site also goes down.

Indeed, if we're to have a public Internet health meter, it must be distributed and hosted/served from "outside" somehow, to be resilient to all or parts of the network being down.

Re: Level 3 Global Outage

#55
Internet infrastructure is broken.

Why do a few companies control the backbone of the internet? Shouldn’t there be a fallback or disaster recovery plan if one or more of these companies become unavailable?

Re: Level 3 Global Outage

#56
post #55

Internet infrastructure is broken. Why do a few companies control the backbone of the internet? Shouldn’t there be a fallback or disaster recovery plan if one or more of these companies become unavailable?

Why doesn't stuff just route around this automatically, if one provider has problems?

Re: Level 3 Global Outage

#57
post #16

I was doing development work which uses a server I've got hosted on digital ocean. I started getting intermittent responses which I thought weird as I hadn't changed anything on the server. I spent a good ten minutes trying to debug the issue before searching for something on duckduckgo, which also didn't respond. Cloudfare shouldn't be involved at all with my little site, so I don't think it's limited to just them.

As noticed in another comment I see loads of problems within Cogentco, all on *.atlas.cogentco.com. Might the problem lies there?

Cogent and Cox are also having problems, but we are seeing a lot more successful traffic on Cogent than CenturyLink. It appears that CL is also not withdrawing stale routes. It seems CLs issues are causing issues on/with everything connected to it.

Re: Level 3 Global Outage

#58

Odd, I'm trying to reach a host in Germany (AS34432) from Sweden but get rerouted Stockholm-Hamburg-Amsterdam-London-Paris-London-Atlanta-São Paulo after which the packets disappear down a black hole. All routing problems occur within Cogentco. 3 sth-cr2.link.netatonce.net (85.195.62.158) 4 te0-2-1-8.rcr51.b038034-0.sto03.atlas.cogentco.com 5 be3530.ccr21.sto03.atlas.cogentco.com (130.117.2.93) 6 be2282.ccr42.ham01.a…

What seems to have happened is that Centurylinks internal routing has collapsed in some way. But they're still announcing all routes and they don't stop announcing routes when other ISPs tag their routes not to be exported by Centurylink.

So as other providers shut down their links to Centurylink to save themselves the outgoing packets towards centurylink travel to some part of the world where links are not shut down yet.

Post reply on HN