Live data from Hacker News

Cloudflare was down

cloudflare.com

351–360 of 560 posts

Re: Cloudflare was down

#352

Wow, just plain 500s on customer sites. That's a level of down you don't see that often.

So. I don't understand the 5 nines they promote. One bad day those nines are gone. So they next year you are pushing 2 nines.

Its just fabricated bullshit. It's how all the companies do it. 99.999% over a year is literally 5 minutes. Or under an hour in a decade, that's wildly unrealistic.

Reddit was once down for a full day and that month they reported 99.5% uptime instead of 99.99% as they normally claimed for most months.

There is this amazing combination of nonsense going on to achieve these kinds of numbers:

1. Straight up fraudulent information on status page. Reporting incendents as more minor than any internal monitors would claim.

2. If it's working for at least a few percent of customers it's not down. Degraded is not counted.

3. If any part of anything is working then it's not down. For example with the reddit example even if the site was dead as long as the image server is still at 1% functional with some internal ping the status is good.

Re: Cloudflare was down

#354
post #23

Earlier quoted context omitted.

hmm... https://downdetectorsdowndetector.com/ (edit: it's working now (detecting downdetector's down))

So, This one is green: https://downdetectorsdowndetector.com This one is not openning: https://downdetectorsdowndetectorsdowndetector.com This one is red: https://downdetectorsdowndetectorsdowndetectorsdowndetector....

it's like they didn't fully think it through/expect people to actually use it so soon

Re: Cloudflare was down

#356
post #126
post #25

This is not good. One major outage? Something exceptional. Several outages in a short time? As someone thats worked in operations, I have empathy; there are so many “temp havks” that are put in place for incidents. but the rest of the world won’t… they’re gonna suffer a massive reputation loss if this goes on as long as the last one.

We are now seeing which companies do not consider the third party risk of single point of failures in systems they do not control as part of their infrastructure and what their contingency plan is. It turns out so far, there isn't one. Other than contacting the CEO of Cloudflare rather than switching on a temporary mitigation measure to ensure minimal downtime. Therefore, many engineers at affected companies would ha…

Alternative infrastructure costs money, and it's hard to get approval from leadership in many cases. I think many know what the ideal solution looks like, but anything linked to budgets is often out of the engineer's hands.

In some cases it is also a valid business decision. If you have 2 hour down time every 5 years, it may not have a significant revenue impact. Most customers think it's too much bother to switch to a competitor anyway, and even if it were simple the competition might not be better. Nobody gets fired for buying IBM

The decision was probably made by someone else who moved on to a different company, so they can blame that person. It's only when down time significantly impacts your future ARR (and bonus) that leadership cares (assuming that someone can even prove that they actually lose customers).

Re: Cloudflare was down

#358
post #214

Earlier quoted context omitted.

Yeah. I only work for a small company, but you can be certain we will not update the status page if only a small portion of customers are affected, and if we are fully down, rest assured there will be no available hands to keep the status page updated

>rest assured there will be no available hands to keep the status page updated That's not how status pages if implemented correctly work. The real reason status pages aren't updated is SLAs. If you agree on a contract to have 99.99% uptime your status page better reflect that or it invalidates many contracts. This is why AWS also lies about it's uptime and status page. These services rarely experience outages accordi…

SLA’s usually just give you a small credit for the exact period of the incident, which is arymetric to the impact. We always have to negotiate for termination rights for failing to meet SLA standards but, in reality, we never exercise them.

Reality is that in an incident, everyone is focused on fixing issue, not updating status pages; automated checks fail or have false positives often too. :/

Re: Cloudflare was down

#359

This is painful, if I'm not mistaken this is during a scheduled maintenance too ? Whenever I deploy a new release to my 5 customers, I am pedantic about having a fast rollback.. Maybe I'm not following the apparent industry standard and instead should just wing it.

Let AI wing it instead.

Re: Cloudflare was down

#360
post #39

It's configuration error or related to configuration. It always is with this big things. Nice thing about Cloudflare being down is that almost everything is down at once. Time for peace and quiet.

Damn, I wish CloudFlare being down also affected local development, so I could take a break from doing frontend… :'(
Post reply on HN