Live data from Hacker News

Today's Outage Post Mortem

blog.cloudflare.com

31–40 of 159 posts

Re: Today's Outage Post Mortem

#31
post #15

Earlier quoted context omitted.

That works when the interfaces are totally standard, but edge/core routers are not like that. Cisco supports one set of protocols for talking to other Cisco products; another set for talking to everything else. The "everything else" protocols suck in a lot of ways (they're ok inter-site, but not really so great intra-site). Same with Juniper. (there aren't really other viable options besides those two) You could buil…

Re: "there aren't really any viable options..." Total misconception. BGP, OSPF, ISIS, LISP, etc. are all non proprietary. Sure, the root cause of this particular problem is that CF is using something specific to Juniper, however router interoperability is not predicated on components like that. This example was a tool CF operationalized, and likely had little to do with their routing with the exception of it being a…

"there aren't really any viable options" to juniper or cisco for core/edge routers.

There are some routing protocols which interoperate (which is how different sites on the Internet can talk to each other), but most of the protocols used for HA or management of a given set of routers, or, more importantly, most tested/debugged implementations of HA and device management, are Cisco or Juniper specific.

No big deal announcing routes to your upstream if you use Juniper and they use Cisco. Big deal if you have Cisco+Juniper and want to do HSRP (Cisco-only).

Re: Today's Outage Post Mortem

#32
post #21

To the couldflare folks; It's refreshing to see you take responsibility, but I think you've been a bit too hard on yourselves by taking all the blame. First of all, what you hit was a unknown bug in JunOS, and Juniper is to blame for their part. Using some form of staging to slow roll-out of rule changes might have saved you from a full meltdown, but when you're getting attacked, every second counts. Slow versus fast…

"by the time I saw the "cloudflare is down" post in the newest queue, it was already back up running again." Not sure what your timeline shows but the "cloudflare is down" post hit the #1 spot on the front page just a few minutes after they went down. About 40 minutes after that, the services started to come back online for me. That's a significant outage. That's not reflecting on the job they did bringing things bac…

Indeed. This was no few minute outage. There was very close to a full hour of zero traffic according to our MRTG.

Re: Today's Outage Post Mortem

#33
That was a pretty interesting writeup and I always like it when companies are totally (and quickly) upfront about negative events.

One thing that occurred to me though is that performing a hard reboot of the routers required calling people to physically access the devices and took some time to perform (as you would expect). Although I wouldn't expect it to be needed very often, I'm sort of surprised CloudFlare doesn't have out-of-band remote power cycle capabilities.

There may be some factor I'm not considering that would make that an unattractive option, but it does seem like it could cut down an already quick response time even further for any similar events in the future.

Re: Today's Outage Post Mortem

#34
post #26

As always I'm glad to see Cloudflare post such detailed outage reports. They are one of the few providers I know of that is willing to go into such depth and that is one of the things I appreciate about them. That said, the outage that occurred was one that was indeed fully preventable. We don't exactly have as many locations as they do, but for internal resources at least, not pushing configuration changes to all de…

Although I agree, it would be rather hard to fight an attack if you didn't roll out fairly quickly...

Re: Today's Outage Post Mortem

#35
post #15
post #9

if you want to build a reliable system, one useful thing to do is use equipment from multiple vendors. sure it's inconvenient, but by doing this you can often de-correlate failures. especially if you want to improve someone else's reliability. e.g., from simple things like hard drives in a raid from different vendors, to n-version programming in safety critical systems (like airplanes).

That works when the interfaces are totally standard, but edge/core routers are not like that. Cisco supports one set of protocols for talking to other Cisco products; another set for talking to everything else. The "everything else" protocols suck in a lot of ways (they're ok inter-site, but not really so great intra-site). Same with Juniper. (there aren't really other viable options besides those two) You could buil…

http://www.zdnet.com/uk-internet-hit-by-linx-router-failure-...

Here's an example of a large provider with parallel infrastructure each powered by a different hardware provider(Brocade/Extreme). One failed and one kept working. I seem to recall a more detailed RFO but my Google-fu has failed me this morning.

Re: Today's Outage Post Mortem

#36

> CloudFlare currently runs 23 data centers worldwide. Shouldn't that always say - CloudFlare currently runs in 23 data centers worldwide? Or is that just how one would phrase that if you rent multiple racks or a cage in a datacenter? ...because I've seen that a bunch of times before from just about everyone. Just curious.

You are correct. However, it is not as bad as the egregious use of the word "unlimited" by hosting providers.

Re: Today's Outage Post Mortem

#37
post #35
post #15

Earlier quoted context omitted.

That works when the interfaces are totally standard, but edge/core routers are not like that. Cisco supports one set of protocols for talking to other Cisco products; another set for talking to everything else. The "everything else" protocols suck in a lot of ways (they're ok inter-site, but not really so great intra-site). Same with Juniper. (there aren't really other viable options besides those two) You could buil…

http://www.zdnet.com/uk-internet-hit-by-linx-router-failure-... Here's an example of a large provider with parallel infrastructure each powered by a different hardware provider(Brocade/Extreme). One failed and one kept working. I seem to recall a more detailed RFO but my Google-fu has failed me this morning.

LINX just runs inter-provider switch fabric, though, which is vastly simpler, and just runs two separate switch fabrics for customers to plug into.

Running an anti-DDoS/CDN service which handles traffic like Cloudflare does would be vastly more difficult.

It's certainly possible to do, but I think the given ~reasonable engineering resources, the net reliability of a heterogenous J/C version of CloudFlare would be less, and performance worse, than what they have now.

Switch fabric is a lot closer to the "run different models of hard drives" (although, you don't do that WITHIN a RAID group either -- you do it on separate RAIDs and possibly separate chassis), than routing infrastructure (which is like running a 777 with 1 GE engine and 1 RR engine. At best, you can turn it back into a 747 and run 2 GE engines and 2 RR engines.)

Re: Today's Outage Post Mortem

#38
post #32

Earlier quoted context omitted.

"by the time I saw the "cloudflare is down" post in the newest queue, it was already back up running again." Not sure what your timeline shows but the "cloudflare is down" post hit the #1 spot on the front page just a few minutes after they went down. About 40 minutes after that, the services started to come back online for me. That's a significant outage. That's not reflecting on the job they did bringing things bac…

Indeed. This was no few minute outage. There was very close to a full hour of zero traffic according to our MRTG.

Pingdom reported an hour downtime on my site too.

Re: Today's Outage Post Mortem

#39
post #37
post #35

Earlier quoted context omitted.

http://www.zdnet.com/uk-internet-hit-by-linx-router-failure-... Here's an example of a large provider with parallel infrastructure each powered by a different hardware provider(Brocade/Extreme). One failed and one kept working. I seem to recall a more detailed RFO but my Google-fu has failed me this morning.

LINX just runs inter-provider switch fabric, though, which is vastly simpler, and just runs two separate switch fabrics for customers to plug into. Running an anti-DDoS/CDN service which handles traffic like Cloudflare does would be vastly more difficult. It's certainly possible to do, but I think the given ~reasonable engineering resources, the net reliability of a heterogenous J/C version of CloudFlare would be les…

I'm not going to go much further because this debate is useless without a context of limits and expectations. No one is discussing simplicity. LINX's operation is not simple. A specialized provider is offloading a difficult function as a core competency in return for simplicity. As an end-user, the difficulty is a non-factor. Just make it happen. Couple in "reasonable" with expectations and then we know what to expect. If it costs the moon to never make this happen again, then charge accordingly. If this happens once in a blue moon, then charge a lesser price.

Re: Today's Outage Post Mortem

#40
> Even though some data centers came back online initially, they fell back over again because all the traffic across our entire network hit them and overloaded their resources.

I know very little of networking, but this seems to be a recurring pattern that aggravates many major outages. What surprises me is that this so often seems to be a scenario not accounted for.

Post reply on HN