Live data from Hacker News

Cloudflare API Down

cloudflarestatus.com

161–170 of 211 posts

Re: Cloudflare API Down

#161
post #150
post #46

When I worked there (3+ years ago), if PDX were out then "the brain" was out... things like DDoS protection was already being done within each PoP (so that will be just fine, even for L3 and L7 floods, even for new and novel attacks), but nearly everything else was done with the compute in PDX and then shipped to each PoP as configuration data. The lifecycle is: PoPs generate/gather data > send to PDX > compute in PD…

What is PDX?

I assume shorthand for a datacenter in Portland. Labeling data centers by their cities airport code is something I see a lot.

Re: Cloudflare API Down

#162
post #150
post #46

When I worked there (3+ years ago), if PDX were out then "the brain" was out... things like DDoS protection was already being done within each PoP (so that will be just fine, even for L3 and L7 floods, even for new and novel attacks), but nearly everything else was done with the compute in PDX and then shipped to each PoP as configuration data. The lifecycle is: PoPs generate/gather data > send to PDX > compute in PD…

What is PDX?

It's the IATA code for Portland International Airport. Many datacenters use IATA codes for the nearest airport to give a rough approximation of the location. So the PDX datacenter is the one closest to the PDX airport in or around Portland.

Re: Cloudflare API Down

#163
post #158
post #71

Earlier quoted context omitted.

Lack of preventive maintenance if I were to guess. Also, these generators would need a supply of diesel fuel, and typically have a storage tank on site. If the diesel isn't used and replaced, it can gum up the generator.

I've gotten 60 year old tractors to run on 60 year old diesel. Gumming up is much more common in gas applications. I guess modern diesel might not be so robust, I know almost nothing about modern engines.

There is nothing so satisfying as when an old engine with bad gas finally catches and starts running continuously.

Re: Cloudflare API Down

#164
post #82

The most impressive thing is: The status page actually works

Isn't it ironic that of all the services failing the status page never seems to go down? Of course it's likely hosted somewhere else, but I always chuckled and thought the same thing when there was a major outage somewhere.

I think that's less of a coincidence than due to the advice many people give that you separate your status page from the infrastructure it's reporting on in every single way possible. One example of the approach is the separate top-level domain usually used. If I were setting a status page up, I would ensure it was with a different cloud provider, and ideally a separate side of the country or separate continent from my principal DC.

Re: Cloudflare API Down

#165
post #105
post #96

Earlier quoted context omitted.

Fun fact, on certain (major) cloud providers, in certain regions, AZs are sometimes different floors of the same building :)

You may be obligated not to name them, but I'm not: Google.

AZ is a term used by AWS and Azure. GCP documentation makes it clear to "Distribute your resources across multiple zones and regions", where regions are physically different data centers.

Re: Cloudflare API Down

#166
post #41

Earlier quoted context omitted.

No free (or even cheap) alternatives exist. If you have a little site that might be a DoS target, you have to use it.

How many little sites do you run that get hit by DDoS? I personally run about 10 tiny websites myself, some of them have around ~1-2K daily active users, but neither of them have suffered from any DDoS frequently nor do they use CloudFlare at all. One has been hit once by a DDoS that kept trying for ~2 days to bring the site down, but a simple "ban IPs based on hitting rate limits" did the trick to avoid going down,…

I wrote this before here, but my site (small b2b saas with a few 100 avid users from small-medium sized companies) gets hit by massive DDOSs a few times a year. The only way I can protect against that is CF bot fight. Everything else will just immediately kill the service until it's over. The last one lasted 24 hours; there were millions of requests from 100000s unique ips over that time; many ips from azure, gcp and aws. Why? I don't have a clue but with CF you simply notice nothing at all.

I cannot rate limit on the machines itself as they die immediately, so then I need to get more advanced firewalls etc which are vastly more expensive than CF.

Re: Cloudflare API Down

#167
post #41
post #37

I dunno. Cloudflare gives me the creeps. I have no idea why so many folks think large swaths of the Internet should be reliant on a single company.

No free (or even cheap) alternatives exist. If you have a little site that might be a DoS target, you have to use it.

Joshua Moon is laughing at you.

Re: Cloudflare API Down

#168
post #50

Earlier quoted context omitted.

Someone else posted about PDX02 going down entirely[0], so sounds like this is the root cause, especially with the latest status update. > Cloudflare is assessing a loss of power impacting data centres while simultaneously failing over services. > [0]: Looks like they lost utility, switched to generator, and then generator failed (not clear on scope of Gen failure yet). Some utility power is back, so recovery is in p…

I think every datacenter I've ever worked with, across ~4 jobs, has had an incident report like "generator failed right as we had an outage." Am I unlucky, or is there something I miss about datacenter administration that makes it really hard to maintain a generator? I guess you don't hear about times the generator worked, but it feels like a high rate of failure to me.

These experiences of power outages is weird to me. What I consider "typical" data center design should make it really hard to lose power.

"Typical" design would be: Each cabinet fed by 2 ATS (transfer switch). Each ATS fed by two UPS (battery bank). Each UPS fed by utility with generator backup. The two ATS can share one UPS/generator, so each cabinet would be fed by 3 UPS+generator. A generator failing to start shouldn't be a huge deal, your cabinet should still have 2 others.

The data center I'm currently in did have a power event ~3 years ago, I forget the exact details but Mistakes Were Made (tm). There were several mistakes that led to it, including that one of the ATS had been in "maintenance mode", because they were having problems getting replacement parts, but then something else happened as well. In short, they had gotten behind on maintenance and no longer had N+1 redundancy.

On top of that, their communication was bad. It was snowing cats and dogs, we suddenly lose all services at that facility (an hour away), and I call and their NOC will only tell me "We will investigate it." Not a "We are investigating multiple service outages", just a "we will get back to you." I'm trying to decide if I need to drive multiple hours in heavy snow to be on site, and they're playing coy...

Re: Cloudflare API Down

#169
post #50

Earlier quoted context omitted.

Someone else posted about PDX02 going down entirely[0], so sounds like this is the root cause, especially with the latest status update. > Cloudflare is assessing a loss of power impacting data centres while simultaneously failing over services. > [0]: Looks like they lost utility, switched to generator, and then generator failed (not clear on scope of Gen failure yet). Some utility power is back, so recovery is in p…

I think every datacenter I've ever worked with, across ~4 jobs, has had an incident report like "generator failed right as we had an outage." Am I unlucky, or is there something I miss about datacenter administration that makes it really hard to maintain a generator? I guess you don't hear about times the generator worked, but it feels like a high rate of failure to me.

Generator needs to go from 0% to almost 100% output within a period of a few seconds, UPS battery is often only a few minutes, long enough to generator to stand-up. There’s a reason why when you put your hand on the cylinder heads for that big diesel they are warm. Much like the theatre “You are only as good as your last rehearsal”.

Re: Cloudflare API Down

#170
post #139

Earlier quoted context omitted.

I feel the same way. What about Akamai, Fastly, or Okta? Maybe Cloudflare gets more attention because their low end plans are accessible to anyone.

Its not just low end plans. Their pricing is basically the only one that feels fair. They don't charge you for bandwidth, unlike others that try to make on it as much as possible, while at the same time having other services also priced significantly higher.

+ (Global) Cloudflare Workers are amazing compared to Google Cloud Functions or other services that are regional, expensive and slow to start.
Post reply on HN