Live data from Hacker News

Today's Outage Post Mortem

blog.cloudflare.com

41–50 of 159 posts

Re: Today's Outage Post Mortem

#41

This is pretty impressive. Keep in mind most of the team is on the west coast so this happened at 1am on a Sunday and they put up a post mortem within hours. Obviously you would prefer it not happen at all, but that is a great response imo.

But this is not impressive:

"Someone from our operations team is monitoring our network 24/7."

"Someone" seems to indicate "1 person". Not "people are monitoring" but "someone". That's it, one person monitors the network? Like the single night guard at the warehouse?

Re: Today's Outage Post Mortem

#42

This is pretty impressive. Keep in mind most of the team is on the west coast so this happened at 1am on a Sunday and they put up a post mortem within hours. Obviously you would prefer it not happen at all, but that is a great response imo.

Everyone who was needed (management, network operations, and the technical support team) were woken up as soon as this happened. The small team who monitor things during the night were quickly calling people. Some folks physically went to the office to help answers phones; others drove to one of our data centers.

"The small team"

Didn't see this when I posted my comment. So there is a team monitoring not a single person?

Re: Today's Outage Post Mortem

#43
post #41

This is pretty impressive. Keep in mind most of the team is on the west coast so this happened at 1am on a Sunday and they put up a post mortem within hours. Obviously you would prefer it not happen at all, but that is a great response imo.

But this is not impressive: "Someone from our operations team is monitoring our network 24/7." "Someone" seems to indicate "1 person". Not "people are monitoring" but "someone". That's it, one person monitors the network? Like the single night guard at the warehouse?

[deleted]

Re: Today's Outage Post Mortem

#44
post #43
post #41

Earlier quoted context omitted.

But this is not impressive: "Someone from our operations team is monitoring our network 24/7." "Someone" seems to indicate "1 person". Not "people are monitoring" but "someone". That's it, one person monitors the network? Like the single night guard at the warehouse?

[deleted]

It's not an issue of whether they can get backup or not.

(And I don't agree with that anyway a night guard can call 911 and get the police pretty quick.)

The issue is whether there is a single person monitoring the network or several or even two. And is the coverage different during "working" hours? And what about the skills of the person monitoring at 1am vs. during the day?

Seems that after something happens people wise up to the weak points. Remembering the case of a single air traffic controller in some towers and after something went wrong there was such shock that only one person was on duty with no backup.

Re: Today's Outage Post Mortem

#45

> Even though some data centers came back online initially, they fell back over again because all the traffic across our entire network hit them and overloaded their resources. I know very little of networking, but this seems to be a recurring pattern that aggravates many major outages. What surprises me is that this so often seems to be a scenario not accounted for.

You can only account for it by having more hardware and then it's possible more of your hardware will fail which puts you right back to where you started.

Re: Today's Outage Post Mortem

#47
post #15

Earlier quoted context omitted.

That works when the interfaces are totally standard, but edge/core routers are not like that. Cisco supports one set of protocols for talking to other Cisco products; another set for talking to everything else. The "everything else" protocols suck in a lot of ways (they're ok inter-site, but not really so great intra-site). Same with Juniper. (there aren't really other viable options besides those two) You could buil…

Re: "there aren't really any viable options..." Total misconception. BGP, OSPF, ISIS, LISP, etc. are all non proprietary. Sure, the root cause of this particular problem is that CF is using something specific to Juniper, however router interoperability is not predicated on components like that. This example was a tool CF operationalized, and likely had little to do with their routing with the exception of it being a…

I view it less as a cost factor and more as a convenience. Developing expertise with juniper and Cisco takes a lot longer. Each router vendor has its own quirks. Even just buolding software to monitor routers is basically a full time job since its a constantly moving target. New bugs are always coming up....

Re: Today's Outage Post Mortem

#48

That was a pretty interesting writeup and I always like it when companies are totally (and quickly) upfront about negative events. One thing that occurred to me though is that performing a hard reboot of the routers required calling people to physically access the devices and took some time to perform (as you would expect). Although I wouldn't expect it to be needed very often, I'm sort of surprised CloudFlare doesn'…

I've never seen remote power cyclers on big routers in major facilities which have on-site remote hands, even when servers all get both IPMI/LOM board cyclers and physical external cyclers. At most, the routers get a serial port connected to a serial port console server or directly to a modem, and/or an admin ethernet network.

I've seen smaller routers, CSU/DSU, etc. type devices in branch offices on cyclers, though.

I think it's mostly that the routers usually have both good OOB management and good watchdog (reboot on freeze) behavior, and that the PSUs in the bigger routers tend to exceed the per-port power limits of most of the external power cyclers.

It may be a good idea, though.

Re: Today's Outage Post Mortem

#49
Just wondering if the source of the large packets were from a [large] range of hosts or maybe a single host?

Ouch if a single host activity took down ~750k websites - whether deliberate and direct or not.

Re: Today's Outage Post Mortem

#50

this is the type of reason why i stopped using cloudflare. there are just too many eggs in one basket. it's as if their entire service becomes a SPOF to your infrastructure.

What is the alternative? Are you saying you can run a service without a DNS provider? You can always have multiple DNS service providers and CloudFlare is probably one of the best ones.
Post reply on HN