Live data from Hacker News

Today's Outage Post Mortem

blog.cloudflare.com

71–80 of 159 posts

Re: Today's Outage Post Mortem

#71

Earlier quoted context omitted.

You would need another network (not just a vlan) to run this as well, if you are going to try to reach it when nothing else is working.

We just hook up a DSL modem to the OOB network or plug it straight into the OOB interface on a core router. You used to do this with actual modems but it's cheap enough to do it with DSL these days, then you're not dependent on any of your own network to access the device in case of failure.

We've been doing this with mikrotik boxes with either wifi or usb gsm modems depending on the what is available in the location.

Re: Today's Outage Post Mortem

#72
post #21

To the couldflare folks; It's refreshing to see you take responsibility, but I think you've been a bit too hard on yourselves by taking all the blame. First of all, what you hit was a unknown bug in JunOS, and Juniper is to blame for their part. Using some form of staging to slow roll-out of rule changes might have saved you from a full meltdown, but when you're getting attacked, every second counts. Slow versus fast…

Is junpier to blame for the bug in their OS? Or is cloudflare to blame for not testing JunOS enough before relying on that OS?

Buck stops with us. We choose the hardware and software that runs on our network. We test and work around thousands of bugs in it. It was up to us to check range limits before applying them. While we'll never be perfect, one of the things I am most proud of with the CloudFlare team is how quickly we do learn from mistakes.

Re: Today's Outage Post Mortem

#73
post #57

Earlier quoted context omitted.

http://en.wikipedia.org/wiki/Jumbogram > An optional feature of IPv6, the jumbo payload option, allows the exchange of packets with payloads of up to one byte less than 4 GiB

Yes but they were still seeing packets bigger than the MTU of Ethernet (or Sonet or whatever other layer 1/2 tech they're connected to the rest of the net with). It doesn't matter what higher level protocols can handle.

They could've been fragmented IPv6 packets. Or it could've been a bug in their profiler.

Re: Today's Outage Post Mortem

#74
post #31

Earlier quoted context omitted.

"there aren't really any viable options" to juniper or cisco for core/edge routers. There are some routing protocols which interoperate (which is how different sites on the Internet can talk to each other), but most of the protocols used for HA or management of a given set of routers, or, more importantly, most tested/debugged implementations of HA and device management, are Cisco or Juniper specific. No big deal ann…

Thats why software defined networking, Openflow etc, are going to take off, as you can get back control of the protocols and what is going on, and avoid the vendor lockin.

> Thats why software defined networking, Openflow etc, are going to take off

I've been hearing this for a decade. It's still not true. I'm not sure why, either.

Re: Today's Outage Post Mortem

#75
post #63
post #12

Earlier quoted context omitted.

The thing which annoyed me the most was losing all DNS. You really need to have the DNS servers in separate infrastructure (ASN, netblock, while anycasted) so there is never a case where both of your DNS are out for a customer domain. The "CNAME" product looks pretty kludgey.

By the same token you (the customer) should not have all your DNS eggs in one basket.

CloudFlare's CDN bits require you to give them DNS delegation of your stuff, last I looked.

Re: Today's Outage Post Mortem

#76
post #59

Earlier quoted context omitted.

Why would you assume this and judge the company on it, based entirely on an offhanded comment? You couldnt have given them the benefit of the doubt long enough to find another of several comments which clearly indicate there was a team of people working on the issue?

The question remains how many people actually are monitoring the network in order to call the first responders. Is it one person, two or five? And my use of "not impressive" was in reply to someone who said "impressive" but more importantly thought it was "impressive" that they put up a post mortem within hours. That's nice but it doesn't answer the question that I had. I stand behind my comment and re ask the questi…

You seem to be looking for "problems" where non exist. There could be 15000000 people monitoring it. It can still go down.

Re: Today's Outage Post Mortem

#77

OT: I want to pitch cloudflare for our CDN needs. Can someone estimate the scale of cloudflare wrt. akamai (current provider), in terms of operations, consumers etc.?

akamai is about 100x the size and probably 200x the price.

Re: Today's Outage Post Mortem

#78
post #21

To the couldflare folks; It's refreshing to see you take responsibility, but I think you've been a bit too hard on yourselves by taking all the blame. First of all, what you hit was a unknown bug in JunOS, and Juniper is to blame for their part. Using some form of staging to slow roll-out of rule changes might have saved you from a full meltdown, but when you're getting attacked, every second counts. Slow versus fast…

Is junpier to blame for the bug in their OS? Or is cloudflare to blame for not testing JunOS enough before relying on that OS?

One would assume paying a company for an OS should be tested via the developers. Juniper SHOULD have tested that route scheme since they sell mission critical architecture.

However you are most likely right, CloudFlare should have test ed it before rolling out.

Re: Today's Outage Post Mortem

#79

> CloudFlare currently runs 23 data centers worldwide. Shouldn't that always say - CloudFlare currently runs in 23 data centers worldwide? Or is that just how one would phrase that if you rent multiple racks or a cage in a datacenter? ...because I've seen that a bunch of times before from just about everyone. Just curious.

The distinction is a bit arbitrary. As a customer you should care that their service is geographically distributed, not whether they own the buildings where the servers are kept.
Post reply on HN