Live data from Hacker News

Cloudflare outage caused by bad software deploy

blog.cloudflare.com

21–30 of 137 posts

Re: Cloudflare outage caused by bad software deploy

#22
post #6

Kinda wonder at this point what findings exist on their Availability SOC 2, assuming they've gotten one. The repeated outages plus the constant malicious advertising by scammy ad providers through cloudflare are slowly turning me off to the service as a potential enterprise customer. Unfortunate too since plenty of superlatively qualified people build great things there (hat tip to Nick Sullivan), but it seems like t…

You can read our SOC3 (public facing SOC2) if you're curious about your availability question: https://www.cloudflare.com/compliance/

There's a lot of good info in there

Re: Cloudflare outage caused by bad software deploy

#25

Kind of funny that it was a regexp.

I'm reminded of:

"You have a problem, and you decide to use a regexp to solve it. Now you have two problems"

Although of course I'm just kidding and I'm sure that a good regexp probably is the right solution for what they're doing in that instance: they have a lot of bright people.

Re: Cloudflare outage caused by bad software deploy

#26
post #16

Nothing like having what should be a world class company falling prey to the same type of screw-ups that plaque 'the local guy maintaining some wordpress site on a shared server'. Separately there is nothing that says that a company like Cloudflare has to air their dirty laundry (as the saying goes). The vast majority of 'customers' really don't care why something happened at all or the reason. All they know is that…

I fail to see the point of your post? Are you arguing that less transparency and information is a good thing?

I doubt any of the people you are talking about "not caring" read this site to begin with.

Re: Cloudflare outage caused by bad software deploy

#28
post #17

At 1402 UTC we understood what was happening and decided to issue a ‘global kill’ on the WAF Managed Rulesets, which instantly dropped CPU back to normal and restored traffic. That occurred at 1409 UTC. So for about 50 minutes, those who relied on the WAF were open to attack?

Isn't open to DDoS better than can't be reached?

Re: Cloudflare outage caused by bad software deploy

#29
post #28
post #17

At 1402 UTC we understood what was happening and decided to issue a ‘global kill’ on the WAF Managed Rulesets, which instantly dropped CPU back to normal and restored traffic. That occurred at 1409 UTC. So for about 50 minutes, those who relied on the WAF were open to attack?

Isn't open to DDoS better than can't be reached?

Depends on the relative costs of the two options?

Re: Cloudflare outage caused by bad software deploy

#30
How to implement a multi-CDN strategy (streamroot.io): https://news.ycombinator.com/item?id=18399523

Etsy implementing multiple CDN (7 years ago, the CDNcontrol project looks abandoned): https://speakerdeck.com/ickymettle/integrating-multiple-cdn-... https://dyn.com/blog/speaking-with-etsy-about-multi-cdns-and...

Basically: you can try to keep a low TTL DNS, but it'll be more DNS traffic, and 5-10% of traffic takes forever to cut over because nobody respects TTL. Worst case you have just as much down time as before, best case most of your traffic is recovered in a few minutes.

Post reply on HN