Live data from Hacker News

The Cloudflare outage might be a good thing

gist.github.com

151–160 of 209 posts

Re: The Cloudflare outage might be a good thing

#151

I don't know how many times I need to say this, but I will die on this hill. Centralized services don't decrease redundancy. They're usually far more redundant than whatever homegrown solution you can come up with. The difference between centralized and homegrown is mostly psychological. We notice the outages of centralized systems more often, as they affect everything at the same time instead of different systems at…

The problem is creating a single point of failure.

There's no doubt a VM in AWS is exponentially more redundant than my VM running on a couple of Intel NUCs in my closet.

The difference is, when I have a major outage, my blog goes down.

When EC2 has a major outage, all of the blogs go down. Along with Wikipedia, Starbucks, and half the internet.

That single point of failure is the issue.

Re: The Cloudflare outage might be a good thing

#152

Earlier quoted context omitted.

If anything, centralisation shields companies using a hyperscaler from criticism. You’ll see downtime no matter where you host. If you self host and go down for a few hours, customers blame you. If you host on AWS and “the internet goes down”, then customers treat it akin to an act of God, like a natural disaster that affects everyone. It’s not great being down for hours, but that will happen regardless. Most compani…

I think this really depends on your industry. If you cannot give a patient life saving dialysis because you don't have a backup generator then you are likely facing some liability. If you cannot give a patient life saving dialysis because your scheduling software is down because of a major outage at a third party and you have no local redundancy then you are in a similar situation. Obviously this depends on your juri…

Yeah I mentioned banking because of what I was familiar with but medical industry is going to be similar.

But they do differ - it’s never ok for a hospital to be unable to dispense care. But it is somewhat ok for one bank to be down. We just assume that people have at least two bank accounts. The problem the banking regulator faces is that when AWS goes down, all banks go down simultaneously. Not terrible for any individual bank, but catastrophic for the country.

And now you see what a juicy target an AWS DC is for an adversary. They go down on their own now, but surely Russia or others are looking at this and thinking “damn, one missile at the right data Center and life in this country grinds to a halt”.

Re: The Cloudflare outage might be a good thing

#153
post #51

Earlier quoted context omitted.

> Is anyone else as confused as I am about how common anti-openness and anti-freedom comments are becoming on HN? In this specific case I don't think it's about being anti-open? It's that a business with only physical presence in one country selling a service that is only accessible physically inside the country.... doesn't.... have any need for selling compressed air to someone who isn't like 15 minutes away from on…

> In this specific case I don't think it's about being anti-open? It's that a business with only physical presence in one country selling a service that is only accessible physically inside the country.... doesn't.... have any need for selling compressed air to someone who isn't like 15 minutes away from one of their gas stations? But that person might be physically further away at the time they want to order somethi…

I guess GP didn't provide enough info, but to me it looked like it was the underlying infra that is networked

That is I'm assuming:

1. Customers are meatspace only, never use any computer interface 2. The network access is for administration only 3. That administration is exclusively in the US

Re: The Cloudflare outage might be a good thing

#154
post #143

Earlier quoted context omitted.

Same idea with the Crowdstrike bug, it seems like it didn't have much of on effect on their customers, certainly not with my company at least, and the stock quickly recovered, in fact doing very well. For me, it looks like nothing changed, no lessons learned.

what do you mean no lesson learned? seems like you haven't been paying attention..there's always a lesson learned

I believe they mean that Crowdstrike learned that they could screw up on this level and keep their customers....

Re: The Cloudflare outage might be a good thing

#155

I don't know how many times I need to say this, but I will die on this hill. Centralized services don't decrease redundancy. They're usually far more redundant than whatever homegrown solution you can come up with. The difference between centralized and homegrown is mostly psychological. We notice the outages of centralized systems more often, as they affect everything at the same time instead of different systems at…

The problem is creating a single point of failure. There's no doubt a VM in AWS is exponentially more redundant than my VM running on a couple of Intel NUCs in my closet. The difference is, when I have a major outage, my blog goes down. When EC2 has a major outage, all of the blogs go down. Along with Wikipedia, Starbucks, and half the internet. That single point of failure is the issue.

Single point of failure means exactly opposite of what you think it means. If my work depends on 5 services to be up, each service would be a single point of failure, and correlation of failure is good for probability that I can do my work.

Re: The Cloudflare outage might be a good thing

#156

Earlier quoted context omitted.

what do you mean no lesson learned? seems like you haven't been paying attention..there's always a lesson learned

I believe they mean that Crowdstrike learned that they could screw up on this level and keep their customers....

That's true of a lot of "Enterprise" software. Microsoft enjoys success from abusing their enterprise customers what seems like daily at this point.

For bigger firms, the reality is that it would probably cost more to switch EDR vendors than the outage itself cost them, and up to that point, CrowdStrike was the industry standard and enjoyed a really good track records and reputation.

Depending on the business, there are long term contracts and early termination fees, there's the need to run your new solution along side the old during migration, there's probably years of telemetry and incident data that you need to keep on the old platform, so even if you switch, you're still paying for CrowdStrike for the retention period. It was one (major) issue over 10+ years.

Just like with CloudFlare, the switching costs are higher than outage cost, unless there was a major outage of that scale multiple times per year.

Re: The Cloudflare outage might be a good thing

#157
post #21

The problem is far more nuanced than the internet simply becoming too centralised. I want to host my gas station network’s air machine infrastructure, and I only want people in the US to be able to access it. That simple task is literally impossible with what we have allowed the internet to become. FWIW I love Cloudflare’s products and make use of a large amount of them, but I can’t advocate for using them in my prof…

> and I only want people in the US to be able to access it. That simple task is literally impossible with what we have allowed the internet to become. Is anyone else as confused as I am about how common anti-openness and anti-freedom comments are becoming on HN? I don’t even understand what this comment wants: Banning VPNs? Walling off the rest of the world from US internet? Strict government identity and citizenship…

> It’s all so foreign and it feels like the vibe shift happened overnight.

The cultural zeitgeist around the internet and technology has changed, unfortunately. But it definitely didn't happen overnight. I've been witnessing it happen slowly over the past 8-10 years, with it accelerating rapidly only in the last 5.

I think it's a combination of special interest groups & nation states running propaganda campaigns, both with bots and real people, and a result of the internet "growing up." Once it became a global, high-stakes platform for finance and commerce, businesses took over, and businesses are historically risk averse. Freedom and openness is no longer a virtue but a liability (for them).

Re: The Cloudflare outage might be a good thing

#158
post #10

It would be a good thing, if it would cause anything to change. It obviously won't. As if a single person reading this post wasn't aware that the Internet is centralized, and couldn't name specifically a few sources of centralization (Cloudflare, AWS, Gmail, Github). As if it's the first time this happens. As if after the last time AWS failed (or the one before that, or one before…) anybody stopped using AWS. As if a…

It’s too few and far between. It’s gonna make some changes if it’s a monthly event. If businesses start to lose connection for 8 hours every month, maybe the bigger ones are going to run for self hosting or at least some capacity of self hosting.

Yeah, agree. But even in case of 8 hour downtime (it's almost 99% SLA) it isn't beneficial for really small firms.

Re: The Cloudflare outage might be a good thing

#159

Earlier quoted context omitted.

The problem is creating a single point of failure. There's no doubt a VM in AWS is exponentially more redundant than my VM running on a couple of Intel NUCs in my closet. The difference is, when I have a major outage, my blog goes down. When EC2 has a major outage, all of the blogs go down. Along with Wikipedia, Starbucks, and half the internet. That single point of failure is the issue.

Single point of failure means exactly opposite of what you think it means. If my work depends on 5 services to be up, each service would be a single point of failure, and correlation of failure is good for probability that I can do my work.

This is a really interesting point, because I could see a situation where your application requires integration with say 10 services. If they all run on AWS, they either all go down or all run together. If they're all self-hosted, there's a good chance that at any time one of the ten is down, and so your service can't run.

Re: The Cloudflare outage might be a good thing

#160
post #26

Earlier quoted context omitted.

> It would be a good thing, if it would cause anything to change. It obviously won't. I agree wholeheartedly. The only change is internal to these organizations (eg: CloudFlare, AWS) Improvements will be made to the relevant systems, and some teams internally will also audit for similar behavior, add tests, and fix some bugs. However, nothing external will change. The cycle of pretending like you are going to impleme…

the root cause is customers refusing to punish these downtime. Checkout how hard customers punish blackouts from the grid - both via wallet, but also via voting/gov't. It's why they are now more reliable. So unless the backbone infrastructure gets the same flak, nothing is going to change. After all, any change is expensive, and the cost of that change needs to be worth it.

> the root cause is customers refusing to punish these downtime.

ok how do I punish cloudflare -- build my own globally-distributed content-delivery network just for myself so that I can be "decentralized"?

Or should I go to one of their even-larger competitors like AWS or GCP?

What exactly do you propose?

Post reply on HN