"Why is PagerDuty calling me before I've had my coffee?" Call declined.
Verizon and a BGP Optimizer Knocked Large Parts of the Internet Offline
121–130 of 291 posts
Re: Verizon and a BGP Optimizer Knocked Large Parts of the Internet Offline
#122> For example, our own IPv4 route 104.20.0.0/20 was turned into 104.20.0.0/21 and 104.20.8.0/21. [...] The prefixes Cloudflare announces are signed for a maximum size of 20. RPKI then indicates any more-specific prefix should not be accepted, no matter what the path is. Did RPKI help reduce the scope of this incident, by stopping propagation of these faulty routes earlier than otherwise? Or did it have no effect in t…
Anecdotal, but: The article notes that AT&T has implemented RPKI, and a client mentioned to me that he wasn't having problems accessing Cloudflare-hosted infrastructure via his AT&T phone. The rest of his employees were having major issues though via the municipal fiber service provider.
Re: Verizon and a BGP Optimizer Knocked Large Parts of the Internet Offline
#123Cloudflare managed to get an in-depth blog post, one which has incident details, points blame to other parties, and makes some really quite aggressive (for corporate blog posts) claims, all during an incident, and they did all that in 8 hours . I'm impressed. At most other similar size companies, this would take 4 days. And in something like Amazon, it would be 2 weeks of approvals, editing, and review before a water…
Re: Verizon and a BGP Optimizer Knocked Large Parts of the Internet Offline
#124> The RPKI framework that we implemented and deployed globally last year is designed to prevent this type of leak. It enables filtering on origin network and prefix size. The prefixes Cloudflare announces are signed for a maximum size of 20. RPKI then indicates any more-specific prefix should not be accepted, no matter what the path is. Does RPKI prevent Cloudflare from announcing additional /22 routes during an inci…
If every network announced all their routes as /24s — the smallest route generally accepted over the public Internet — the routing table would be a giant mess and would overwhelm many routers' ability to store them.
That said, after today we are thinking about ways that, in case of an emergency, we could break the routes down to be more specific than whatever is leaking. Given how broadly peered we are, Cloudflare's network will be as protected as anyone's. However, that's not really a good solution for the Internet generally. Better that we all implement and enforce RPKI.
Re: Verizon and a BGP Optimizer Knocked Large Parts of the Internet Offline
#125One would think Cloudflare team would have a direct line of communication to all tier 1 Internet providers.
We thought we did. And tried both public and private lines of communication — without reply. Still waiting.
Re: Verizon and a BGP Optimizer Knocked Large Parts of the Internet Offline
#126>"It doesn't cost a provider like Verizon anything to have such limits in place. And there's no good reason, other than sloppiness or laziness, that they wouldn't have such limits in place."
Is "sloppiness or laziness" really the only possible attribution here? I'm not a big fan of Verizon but I'm a big fan of civility and empathy, two qualities which your blog post lacks. Outages are a really unfortunate fact of life. We've seen them recently with Google, AWS, Dyn - all companies where technical competency is generally not questioned. It's quite possible the cause of of this outage was some "perfect storm" scenario such as an eBGP router rebooted and came up with a stale or incorrect config. "Perfect storm" scenarios even happen at companies with very rigorous engineering cultures as we saw with the most recent Google outage.
Your attempt to shame an organization without knowing all the details reeks of immaturity and pettiness. Ditto for your willingness to turn this into yet another Cloudflare marketing opportunity. Have you forgotten about your own Cloudbleed incident? How would you feel if it a security company took that as an opportunity to shame you for "sloppiness or laziness"? Or some other company's CEO was offering to send people "Cloudbleed Support Group" T-Shirts on HN as your own CEO is doing in this thread?
Lastly RPKI isn't a silver bullet, RPKI authorities can also be misconfigured and attacked[1][2]. This happened with the LACNIC incident in 2013[2]. It's also worth mentioning that RPKI potentially creates new threats[2]. But again it seems more important to you to use this as a marketing opportunity and promote yourself while throwing someone else under a bus while uttering pithy summations.
Also from your post:
>"And, in particular, we're looking at you Verizon — and still waiting on your reply."
Although Verizon is the 400lb gorilla in the room, their NOC and network engineers are still regular people with kids and families and feelings. They are also people who have had a really shit day today. Why you can't extend just a bit of human compassion and feel compelled to try to shame is quite inexplicable.
You may think that your blog post was a marketing coup but I see it as a massive failure in in both leadership and civility.
As a thought exercise maybe Cloudflare leadership could think about how they would like the community to react the next time they are at fault.
Re: Verizon and a BGP Optimizer Knocked Large Parts of the Internet Offline
#127I work from home in NJ and I knew something was screwy this morning. I wish there was a place I could have checked to see it was this issue. I rebooted pretty much everything in my house.
Re: Verizon and a BGP Optimizer Knocked Large Parts of the Internet Offline
#128Earlier quoted context omitted.
So we run into the age-old problem of "who decides". Also, how do we prevent fragmentation when there is disagreement.
Freedom isn't free. Web of trust. Inconvenient, but that's a price I'm willing to pay for a network that empowers users rather than commercial interests.
Re: Verizon and a BGP Optimizer Knocked Large Parts of the Internet Offline
#129Earlier quoted context omitted.
The amount of posturing and blaming in Cloudflare's response is breathtakingly unprofessional. If the article was just a few sentences longer, you could have squeezed in a few more statements of blame. We know, they messed up. But Cloudflare isn't making itself look any better by rolling the bus over Verizon again and again.
I 100% agree with you, though I am unsurprised that HN is downvoting you. HN seems to revere Cloudflare bigtime despite the fact that Cloudflare often uses HN as their own corporate PR platform. I absolutely loathe Verizon and I'll be the first to line up for a good publish lashing of US ISPs, but even I feel like this blog post is unnecessarily unprofessional. What strikes me the most is that this whole "event" woul…
1. They implement basic precautions to prevent dumb things from going wrong.
2. They're available 24/7, to immediately respond to and remediate whatever does go wrong.
3. Both of the above are core obligations, which supersede any questions of public relations or maturity or higher-ups not wanting to be bothered.
If Verizon can't be trusted to properly operate their network, that's an immediate threat to the health of the Internet, and many people do need to be made aware of it. It's not just Cloudflare being salty because their customers yelled at them.
Re: Verizon and a BGP Optimizer Knocked Large Parts of the Internet Offline
#130Earlier quoted context omitted.
The amount of posturing and blaming in Cloudflare's response is breathtakingly unprofessional. If the article was just a few sentences longer, you could have squeezed in a few more statements of blame. We know, they messed up. But Cloudflare isn't making itself look any better by rolling the bus over Verizon again and again.
I 100% agree with you, though I am unsurprised that HN is downvoting you. HN seems to revere Cloudflare bigtime despite the fact that Cloudflare often uses HN as their own corporate PR platform. I absolutely loathe Verizon and I'll be the first to line up for a good publish lashing of US ISPs, but even I feel like this blog post is unnecessarily unprofessional. What strikes me the most is that this whole "event" woul…
It wasn't just CloudFlare who were affected. And the time of day is completely irrelevant, I live in Australia and was affected by this during evening peak time. Some very popular services (eg: discord) were completely knocked offline.
I think you're underestimating the impact of this event.