Live data from Hacker News

Cloudflare Network Performance Issues

cloudflarestatus.com

271–280 of 329 posts

Re: Cloudflare Network Performance Issues

#271

Cloudflare: Your status page showed "all systems operational" for over 20 minutes while your primary domain was returning a 502 error. Please change this to update automatically, many other engineering teams depend on you. https://i.imgur.com/qHBM2JW.png

Sadly this reminds me of AWS outages too where the same applies. How is it that hundreds of developers know there's an issue before AWS do, or Cloudflare in this instance. See my blog post on similar AWS uptime reporting issues at https://www.ably.io/blog/honest-status-reporting-aws-service . At Ably, our status site had an incident update about Cloudflare issues being worked on (by routing away from CF) before Cloud…

>How is it that hundreds of developers know there's an issue before AWS do

Trust me, they know.

They know about problems we never find out about, too.

Re: Cloudflare Network Performance Issues

#272

Earlier quoted context omitted.

I guess you’ve never worked in enterprise. SLAs for critical systems very frequently incur payback in excess of the billing on outages.

But not consequential losses which is what the parent mentions.

I'm definitely not suggesting CF should cover losses. Sorry if I gave that impression. That would effectively require them to be an insurance company since they'd have to investigate claims, and possibly charge customers differently based on risk. (i.e. you don't want to bill a customer $200 per month if 10 minutes of downtime could lose $20 million in sales.)

I mean something like Amazon EC2's SLA (https://aws.amazon.com/compute/sla/) where credits are proportional to downtime, but not 1:1. i.e. they credit 100% for >= 5% downtime. With Cloudflare's SLA, 5% downtime (1.5 days in a month) would only give you a 5% credit.

Re: Cloudflare Network Performance Issues

#273
post #255

Once cloudflare.com came back I decided to check out their business SLA, and it's not very encouraging: > For any and each Outage Period during a monthly billing period the Company will provide as a Service Credit an amount calculated as follows: Service Credit = (Outage Period minutes * Affected Customer Ratio) ÷ Scheduled Availability minutes - https://www.cloudflare.com/business-sla/ So assuming an outage affects…

They reserve the best SLA for Enterprise, naturally. https://www.cloudflare.com/plans/enterprise/ > 100% uptime and 25x Enterprise SLA > In the rare event of downtime, Enterprise customers receive a 25x credit against the monthly fee, in proportion to the respective disruption and affected customer ratio.

Two years missing revenue for enterprise customers, ouch that's going to hurt.

Re: Cloudflare Network Performance Issues

#274
post #95

A great example for why you shouldn't transfer your domain to Cloudflare Registrar if you're also using their CDN. Those who have transferred their domains cannot change DNS servers to mitigate the outage.

Changing your NS records at the registry could help, but keep in mind most TLDs are serving NS records with 1-2 day TTLs, so you'll still see a lot of traffic going to the old server.

If this is something you want to be able to mitigate, you really need to be running a seperate DNS infra from your hosting/CDN and use short TTL cnames to delegate hostnames to the CDN. This becomes a big challenge if you host on an apex domain (eg example.org instead of www.example.org), so don't do that.

Re: Cloudflare Network Performance Issues

#275

Earlier quoted context omitted.

Sadly this reminds me of AWS outages too where the same applies. How is it that hundreds of developers know there's an issue before AWS do, or Cloudflare in this instance. See my blog post on similar AWS uptime reporting issues at https://www.ably.io/blog/honest-status-reporting-aws-service . At Ably, our status site had an incident update about Cloudflare issues being worked on (by routing away from CF) before Cloud…

slightly off-topic: Nice and clean, yet detailed enough status page. I like it. :)

Thanks :)

Re: Cloudflare Network Performance Issues

#276

Earlier quoted context omitted.

I get your point. But I'm not sure I can ever grok the mindset of someone who thinks, "good, that'll teach 'em". Not sure I'd ever want to work with that individual.

This is important to recognize. Victim blaming is a highly destructive practice.

Assigning blame for technical business decisions to the people who made those decisions is victim blaming?

Re: Cloudflare Network Performance Issues

#277
post #240

Earlier quoted context omitted.

They came off that way in this blog post: https://blog.cloudflare.com/how-verizon-and-a-bgp-optimizer-...

That didn't seem like trolling – just a public call for Verizon to follow internet best practices. Given that most large ISPs treat failures as a PR exercise, that's probably necessary.

Agreed.

They reached out to Verizon privately, a Tier 1 carrier with expectations and responsibilities as a good netizen, and got no response.

They attempted to reach out through Verizon's public forms of communication and got a bullshit irrelevant CS response despite requesting escalation.

They then called out Verizon before the community as a whole.

They don't have the luxury of waiting for a well prepared letter from some Verizon lawyers. Modern day customer expectations don't allow for it. You may call it trolling, but all I saw was a company asking another company to stop pissing in the public pool.

Re: Cloudflare Network Performance Issues

#278

Once cloudflare.com came back I decided to check out their business SLA, and it's not very encouraging: > For any and each Outage Period during a monthly billing period the Company will provide as a Service Credit an amount calculated as follows: Service Credit = (Outage Period minutes * Affected Customer Ratio) ÷ Scheduled Availability minutes - https://www.cloudflare.com/business-sla/ So assuming an outage affects…

AWS CloudFront SLA refund/credit policy appears to work as you describe: https://aws.amazon.com/cloudfront/sla/

Monthly Uptime:

   * 99.0% 
But CloudFlare's Enterprise SLA (25x credit) is similar or maybe even a little bit better (because you get to 100% at 96% instead of 95%). Of course when you are doing an Enterprise deal you can negotiate for whatever terms are mutually acceptable as long as you're willing to pay.

In any case, the function of the credit policy is to ensure there is enough pain for the provider to put in place the quality / reliability practices, process and code to protect themselves from losses. IMO most sustainable business pull in much more revenue per hour than 25x their CDN cost.

It would be interesting to know how CloudFlare's infra and processes differentiate free, Business and Enterprise customers.

Re: Cloudflare Network Performance Issues

#279
This is just my personal guess, but it's likely China flexing again after the protest escalation yesterday in Hong Kong. A few weeks ago Telegram was attacked by China [0], and Hong Kong protesters used Telegram to communicate.

This time when CloudFlare was down, the most popular local forum among protesters, lihkg.com, was brought down as well.

[0]: https://techcrunch.com/2019/06/12/telegram-faces-ddos-attack...

Re: Cloudflare Network Performance Issues

#280
They have now released an initial statement [1]:

For about 30 minutes today, visitors to Cloudflare sites received 502 errors caused by a massive spike in CPU utilization on our network. This CPU spike was caused by a bad software deploy that was rolled back. Once rolled back the service returned to normal operation and all domains using Cloudflare returned to normal traffic levels.

This was not an attack (as some have speculated) and we are incredibly sorry that this incident occurred. Internal teams are meeting as I write performing a full post-mortem to understand how this occurred and how we prevent this from ever occurring again.

[1] https://blog.cloudflare.com/cloudflare-outage/

Post reply on HN