Live data from Hacker News

Cloudflare had a partial outage

cloudflare.com

431–440 of 448 posts

Re: Cloudflare had a partial outage

#431
post #415
post #138

Earlier quoted context omitted.

Yes https://www.cloudflare.com/case-studies/digitalocean/

Moved all my domains to DO specifically to stop donating traffic data to Cloudflare. Absolutely stunned I didn't notice this earlier. It's bullshit all the way down. Are there any companies left offering free DNS usable from Terraform that aren't part of the "Internet Five Eyes"? edit: looks like Linode may be the next best 'not terrible' option

I have bad news for you, Linode's authoritative DNS service also uses Cloudflare DNS Firewall.

  $ dig +short ns1.digitalocean.com aaaa
  2400:cb00:2049:1::adf5:3a33
  $ dig +short ns1.linode.com aaaa
  2400:cb00:2049:1::a29f:1a63

Re: Cloudflare had a partial outage

#433
post #318

Earlier quoted context omitted.

Never market during another companies outage, offer help instead. Tomorrow it will be you.

I absolutely agree and very respectfully so. No one is immune to outages. Well said. CDNReserve is designed in a way, that if the outage occurs on one platform it will map the traffic to failover CDN and if the failover/backup CDN suffers an outage, the traffic will be shifted to the primary CDN using CDNReserve. Its built on the premise that the likelihood of two CDNs having outage at the same time is close to ZERO.

The likelihood of CDNReserve having an outage on the other hand is 100%.

You aren't the first to come up with the idea of a CDN traffic director (I built one), and you'll soon discover customers recognize you are just another single point of failure and not the solution. Best to focus on the things other companies in the space market on, bill optimization, latency optimization, etc.

Re: Cloudflare had a partial outage

#434
post #95

Earlier quoted context omitted.

To be fair this is not relying on a single SaaS for everything but many people relying on a single SaaS. I mean if you want to use a reverse proxy/CDN, you must rely on someone .

My company uses 3 CDNs, although not cloudflare. If one (say Aakami) goes down it gets removed from the pool and life continues.

what are those 3 CDN companies? Cloudflare, Aakami and what else?

Re: Cloudflare had a partial outage

#435
post #426

Earlier quoted context omitted.

/s means "end of sarcasm".

Ah. Obviously. So yeah, how is a CDN centralizing your infra? You could just have your CNAMEs point to a different provider or directly to your gateways. Or you could even go down the multi CDN path, and have someone like ns1 automatically redirect your CNAMEs to an alternate CDN on a per-geo basis to overcome local failures. It's just another SaaS component in your system. You could self host if you're willing to ta…

Not the OP, but CloudFlare is not only a CDN but does everything you mentioned in your comment for you, so it's the load balancer and the DNS as well. When it goes down everything goes down.

Technically you could set up a separate DNS/failover somewhere else and use a backup reverse proxy/TLS terminator/CDN SaaS similar to CloudFlare, but then that somewhere else will be your point of failure.

Re: Cloudflare had a partial outage

#436

Earlier quoted context omitted.

Soooo... Cloudflare?

No, we're talking about a colocation provider, or a leased dedicated server provider. I went with OVHcloud US for my latest deployment. HN is at m5hosting.com.

OVH had some server fires that caused some amount of user downtime. I'm not really sure how that's gonna help.

Unless you have fallback with multi cloud deployments.

Re: Cloudflare had a partial outage

#437

Earlier quoted context omitted.

But the infrastructures those protocols provide (the IPFS network, torrent swarms) can be an alternative to Cloudflare. Which is why I brought it up

Not to state the obvious, but... if a big centralized company built a Cloudflare for IPFS to make it easy for the masses to adopt, that company could go down just as easy as Cloudflare.

How so? Somebody links to a webpage, decentralized resolver converts it to an IPFS hash, which the client queries for any providers of that hash, and retrieves directly from them. No central authority necessary

Re: Cloudflare had a partial outage

#438

Earlier quoted context omitted.

Our key customer facing services are a 99.995% uptime (and a total of 2 or fewer incidents per year of "any length"), which means once you start concatenating services with 99.995% SLAs you aren't there. How that SLA measures a 2 second outage for some customers is a separate thing, and sort of shows how meaningless these things can be on the internet (if you lose service for 10% of your potential customers is that a…

Measuring outages doesn't seem so meaningless as long as you money seems inaccessible. Their main site went down for about 20 hours a couple weeks ago because their hosting provider went down. They deployed an HTTPS only static site in its stead, so at first blush it looked like they deployed nothing. Great when you're trying to find contact information hosted on that site. Their online banking site leveraged Cloudfl…

Sure but that's a total outage for a long time.

What if for some reason a single /24 was unreachable from the site (say an errant route for 12.85.25.0/24 somehow got in the path). How would you even know that was a problem - how many customers are on that /24, how would I measure their failed attempts to connect?

I have a remote office in India on Tata. The other day it had access to much of the internet, but due to a fibre break in the Mederteranian it didn't have access to end points in Europe for a good 20 seconds.

However the other link on a different ISP remained working at that time.

Does that count as an outage? If I wasn't actively monitoring that link with a high resolution would I even know about it?

Re: Cloudflare had a partial outage

#439

Earlier quoted context omitted.

Measuring outages doesn't seem so meaningless as long as you money seems inaccessible. Their main site went down for about 20 hours a couple weeks ago because their hosting provider went down. They deployed an HTTPS only static site in its stead, so at first blush it looked like they deployed nothing. Great when you're trying to find contact information hosted on that site. Their online banking site leveraged Cloudfl…

Sure but that's a total outage for a long time. What if for some reason a single /24 was unreachable from the site (say an errant route for 12.85.25.0/24 somehow got in the path). How would you even know that was a problem - how many customers are on that /24, how would I measure their failed attempts to connect? I have a remote office in India on Tata. The other day it had access to much of the internet, but due to…

I'd argue you're starting from a few orders of magnitude more competency than the credit union was. Their non-banking site was hosted by some podunk company in Texas with no sense of redundancy anywhere. Their provider had a near total networking outage and the credit union had no plan to recover from that.

Insofar as proactively monitoring a single /24, you (probably) don't. I don't think it's (usually) a company's job to monitor their customer's ISPs. The failures that "my" credit union had were due to their own choice in infra (Armor, Cloudflare). When Sonic nuked my config on their DSLAM after some maintenance I raised an issue with Sonic not with whatever other companies became inaccessible as a result.

> Does that count as an outage?

My POV may very well differ from whatever contracts and SLAs you have in place, but yeah maybe. If you can't fail over to the alternative ISP then yes that's an outage. Of course a trans-atlantic fiber break would also likely be a lot more noticeable than fat fingering a route for a /24. And sure, I've been stuck at megacorp when the VPN started handing out addresses in a new subnet but our department's networking team hadn't caught up. That's why you listen to your customers instead of throwing out a "someone else screwed up there's nothing we can do" response.

Me personally I don't think that a 20 minute banking outage is a massive problem (I've long since moved my money elsewhere), even the 20 hour outage was relatively minor. It just speaks to the unwillingness of the credit union to be highly available. They knew of the Armor outage and didn't actually test the remediation. I assume they didn't know about the Cloudflare outage. Both worry me. What happens when they're faced with a total failure of their online banking system?

Re: Cloudflare had a partial outage

#440

Earlier quoted context omitted.

My company uses 3 CDNs, although not cloudflare. If one (say Aakami) goes down it gets removed from the pool and life continues.

what are those 3 CDN companies? Cloudflare, Aakami and what else?

> what are those 3 CDN companies? Cloudflare, Aakami and what else?

The OP says "not Cloudflare"; so probably Akamai, Fastly, CloudFront?

Post reply on HN