Live data from Hacker News

Cloudflare outage on June 21, 2022

blog.cloudflare.com

71–80 of 234 posts

Re: Cloudflare outage on June 21, 2022

#71
post #63

If I use Cloudflare, what can I do — if anything — to avoid disruption when they go down?

On the enterprise plans, you are able to set up your own DNS server that can route users away from Cloudflare, either to your origin or to another CDN/proxy.

Re: Cloudflare outage on June 21, 2022

#73

Yet another BGP caused outage. At some point we should collect all of them: - Cloudflare 2022 (this one) - Facebook 2021: https://news.ycombinator.com/item?id=28752131 - this one probably had the single biggest impact, since engineers got locked out of their systems, which made the fixing part look like a sci-fi movie - (Indirectly caused by BGP: Cloudflare 2020: https://blog.cloudflare.com/cloudflare-outage-on-july-…

These are the public facing BGP announcements that cause problems, but doesn't account for the ones on private LANs that also happen. Previous employers of mine have had significant internal network issues because internal BGP between sites started causing problems. I'm not sure there's anything better (I am not a network guy), but this list can't be exhaustive.

Re: Cloudflare outage on June 21, 2022

#74
post #68

Yet another BGP caused outage. At some point we should collect all of them: - Cloudflare 2022 (this one) - Facebook 2021: https://news.ycombinator.com/item?id=28752131 - this one probably had the single biggest impact, since engineers got locked out of their systems, which made the fixing part look like a sci-fi movie - (Indirectly caused by BGP: Cloudflare 2020: https://blog.cloudflare.com/cloudflare-outage-on-july-…

The internet runs on BGP, I would think that most internet issues would be a result of BGP then.

There are lots of other causes of incidents, like cut cables, failed router hardware, data centers losing power etc.

It just seems that most of these are local enough and the Internet resilient enough that they don't cause global issues. Maybe the exception would be AWS us-east-1 outages :-)

Re: Cloudflare outage on June 21, 2022

#75
One of our sites uses Cloudflare and serves 400k pageviews per month and generates around $650/day in ad and affiliate revenue. If the site is not up the business is not making any money.

Looking at the hourly chart in Google Analytics (compared to the previous day) there isn't even a blip during this outage.

So for all the advantages we get from Cloudflare (caching, WAF, security [our WP admin is secured with Cloudflare Teams], redirects, page rules, etc) I'll take these minor outages that make HN go apeshit.

Of course it helped that most our traffic is from the US and this happened when it did but in the past week alone we served over 180 countries which Cloudflare helps make sure is nice and fast :D

Re: Cloudflare outage on June 21, 2022

#77

In a world where it can take weeks for other companies to publish a postmortem after an outage (if they ever do), I never ceases to amaze me how quickly CF manage to get something like this out. I think it's a testament to their Ops/Incident response teams and internal processes, it builds confidence in their ability to respond quickly when something does go wrong. Incredible work!

I agree, I think the transparency builds trust and I encourage it where I can. The counter thought I had when reading this case though, is it almost feels too fast. What I mean by that is I hope there isn't an incentive to wrap up the internal investigation quickly and write the blog and send it, and go we're done. Doing incident response (both outage and security), the tactical fixes for a specific problem are usual…

I have to agree. The environment that leads to a fast blog post may also lead to this quote from the post:

> This was delayed as network engineers walked over each other's changes, reverting the previous reverts, causing the problem to re-appear sporadically.

They are running as fast as they can and this extended the incident. There is a “slow is smooth, smooth is fast” lesson in here. I’d rather have a team that takes a day to put up the blog post, but doesn’t unnecessarily extend downtime because they are sprinting.

Re: Cloudflare outage on June 21, 2022

#79

Earlier quoted context omitted.

I feel like others lose opportunities by not doing the same. By publishing early and publishing the details they: keep the company in the news with positive stuff (free ad), get an internal documentation of the incident (ignoring the customer oriented "we're sorry" part), effectively get a free recruitment post (you're reading this because you're in tech and we do cool stuff, wink), release some internal architecture…

> I wonder how much those posts are calculated and how much organic/culture related. Don't companies have a fiduciary duty to calculate things; the reason for doing something actually cannot just be that it's a nice thing to do? Not down to the word, but at least the general decision to be this way?

No they don't have such duty. In practice very little decision making is based on hard data in my experience. Real world being fuzzy and risk being hard to quantify do not help the situation.

Re: Cloudflare outage on June 21, 2022

#80

Earlier quoted context omitted.

To further add to your point, the CTO is the one who shared it here & the CEO is incredibly active on forums & social media everywhere with customers. Communication has always been one of their strengths.

I do wonder what would happen should happen if either of them left the company, I feel like there's a lot of trust on HN (and other places) that's heavily attached to them as individuals and their track record of good communication.

It could be good or bad; I suspect they've thought about it and have worked on succession (I hope!) and have like-minded people in the wings.

But once it happens things will change and, to be honest, likely for the worse.

edit> fix typo

Post reply on HN