Earlier quoted context omitted.
You would need another network (not just a vlan) to run this as well, if you are going to try to reach it when nothing else is working.
We just hook up a DSL modem to the OOB network or plug it straight into the OOB interface on a core router. You used to do this with actual modems but it's cheap enough to do it with DSL these days, then you're not dependent on any of your own network to access the device in case of failure.
Today's Outage Post Mortem
71–80 of 159 posts
Re: Today's Outage Post Mortem
#72To the couldflare folks; It's refreshing to see you take responsibility, but I think you've been a bit too hard on yourselves by taking all the blame. First of all, what you hit was a unknown bug in JunOS, and Juniper is to blame for their part. Using some form of staging to slow roll-out of rule changes might have saved you from a full meltdown, but when you're getting attacked, every second counts. Slow versus fast…
Is junpier to blame for the bug in their OS? Or is cloudflare to blame for not testing JunOS enough before relying on that OS?
Re: Today's Outage Post Mortem
#73Earlier quoted context omitted.
http://en.wikipedia.org/wiki/Jumbogram > An optional feature of IPv6, the jumbo payload option, allows the exchange of packets with payloads of up to one byte less than 4 GiB
Yes but they were still seeing packets bigger than the MTU of Ethernet (or Sonet or whatever other layer 1/2 tech they're connected to the rest of the net with). It doesn't matter what higher level protocols can handle.
Re: Today's Outage Post Mortem
#74Earlier quoted context omitted.
"there aren't really any viable options" to juniper or cisco for core/edge routers. There are some routing protocols which interoperate (which is how different sites on the Internet can talk to each other), but most of the protocols used for HA or management of a given set of routers, or, more importantly, most tested/debugged implementations of HA and device management, are Cisco or Juniper specific. No big deal ann…
Thats why software defined networking, Openflow etc, are going to take off, as you can get back control of the protocols and what is going on, and avoid the vendor lockin.
I've been hearing this for a decade. It's still not true. I'm not sure why, either.
Re: Today's Outage Post Mortem
#75Earlier quoted context omitted.
The thing which annoyed me the most was losing all DNS. You really need to have the DNS servers in separate infrastructure (ASN, netblock, while anycasted) so there is never a case where both of your DNS are out for a customer domain. The "CNAME" product looks pretty kludgey.
By the same token you (the customer) should not have all your DNS eggs in one basket.
Re: Today's Outage Post Mortem
#76Earlier quoted context omitted.
Why would you assume this and judge the company on it, based entirely on an offhanded comment? You couldnt have given them the benefit of the doubt long enough to find another of several comments which clearly indicate there was a team of people working on the issue?
The question remains how many people actually are monitoring the network in order to call the first responders. Is it one person, two or five? And my use of "not impressive" was in reply to someone who said "impressive" but more importantly thought it was "impressive" that they put up a post mortem within hours. That's nice but it doesn't answer the question that I had. I stand behind my comment and re ask the questi…
Re: Today's Outage Post Mortem
#77OT: I want to pitch cloudflare for our CDN needs. Can someone estimate the scale of cloudflare wrt. akamai (current provider), in terms of operations, consumers etc.?
Re: Today's Outage Post Mortem
#78To the couldflare folks; It's refreshing to see you take responsibility, but I think you've been a bit too hard on yourselves by taking all the blame. First of all, what you hit was a unknown bug in JunOS, and Juniper is to blame for their part. Using some form of staging to slow roll-out of rule changes might have saved you from a full meltdown, but when you're getting attacked, every second counts. Slow versus fast…
Is junpier to blame for the bug in their OS? Or is cloudflare to blame for not testing JunOS enough before relying on that OS?
However you are most likely right, CloudFlare should have test ed it before rolling out.
Re: Today's Outage Post Mortem
#79> CloudFlare currently runs 23 data centers worldwide. Shouldn't that always say - CloudFlare currently runs in 23 data centers worldwide? Or is that just how one would phrase that if you rent multiple racks or a cage in a datacenter? ...because I've seen that a bunch of times before from just about everyone. Just curious.
Re: Today's Outage Post Mortem
#80I think you should have investigated why you got ~90kb packages despite having a max pkg size of ~4kb instead of putting in that rule. :)