Live data from Hacker News

Salesforce Global Outage

status.salesforce.com

151–160 of 184 posts

Re: Salesforce Global Outage

#151
post #66

Despite all of the snark here, in my experience Salesforce SRE team is quite competent. The engineering challenges of running a large PaaS - not just with own apps, but with millions of customer-written apps running on it - are quite interesting, and sadly things happen. The status page makes sense to actual customers, it's the particular "pods" where a given service runs.

Hacker News is much easier to read when you realize that 95% of people have never worked on a "high" (maybe we could say >1B requests per day as a starting point) scale distributed service and think it's trivial to run one with more than 2 nines. You see comments all the time here mentioning that their own desktop at home is achieving more than that which belies deep misunderstanding of how systems are measured. Or t…

[deleted]

Re: Salesforce Global Outage

#152
post #149

Earlier quoted context omitted.

While not a home-run server, the NTP system is a distributed service that receives 100 billion to trillions of requests per day, and it's running pretty smoothly - it's never gone down completely since it started in 1985. It's also very simple. The reason it has so many 9's uptime is because it is simple. Given a low amount of complexity, it's not unreasonable to think that an individual could run a >1B requests per…

>Salesforce is not simple. It's wildly, overly complex. It turns out business environments are wildly overly complex.

I think it’s like advertising - 50% of my code is wildly over complicated - I just don’t know which 50%

But the GP is essentially correct - there is a 2% of salesforce that could be built run and keep 80% of salesforce users happy. Except that you could not charge enough to be able to advertise on F1 cars and take SVPs out to dinner.

So you could not actually make 80% of them happy - they would ever buy it.

Re: Salesforce Global Outage

#153

Earlier quoted context omitted.

In my entire adult lifetime (mid 40s), my ISP has never been out for five days. Compare to Github, Microsoft, Salesforce, and AWS outages that are always occurring in some fashion. Reddit is down constantly in various ways and still continues to operate as a business, public no less, so I disagree about the need to chase five nines and broadly speaking, large distributed systems that are potentially unnecessary for t…

Consider yourself lucky that you’ve never been the victim of a fiber cut. But what about if the power to your house goes out? Or what if your server blows the power supply? My entire point is that you have no redundancy in your system and you also aren’t big enough to have any pull with the vendors who can fix these types of outages so you’re basically at the mercy of your providers with no recourse. That’s why these…

I mean, if you really need redundancy, isn’t a second instance on a VPS somewhere that you manually switch over to, enough?

Over multiple ISPs, so far internet outages for more then a few minutes is very rare (though the few minutes would make me not want to host something requiring high availability; and a cut cable is really annoying because there simply is no quick fix), power outages even rares, I experienced 3 in 40 years, and the longest was 6 hours.

Re: Salesforce Global Outage

#154
post #66

Despite all of the snark here, in my experience Salesforce SRE team is quite competent. The engineering challenges of running a large PaaS - not just with own apps, but with millions of customer-written apps running on it - are quite interesting, and sadly things happen. The status page makes sense to actual customers, it's the particular "pods" where a given service runs.

For me personally the snark isn't because their SRE team is incompotent. It's because software that tries to be everything to everyone is inherently terrible. It's not fun to use for the users, and so many compromises need to be made on the technology side to make that happen that it ends up just being crap all around. This includes Salesforce, SAP, Dynamics, any platforms like that which scale many industries.

Flexibility and abstraction come at a high cost. It doesn't really matter though, world domination at all costs is the name of the game.

Re: Salesforce Global Outage

#155
post #66

Despite all of the snark here, in my experience Salesforce SRE team is quite competent. The engineering challenges of running a large PaaS - not just with own apps, but with millions of customer-written apps running on it - are quite interesting, and sadly things happen. The status page makes sense to actual customers, it's the particular "pods" where a given service runs.

Hacker News is much easier to read when you realize that 95% of people have never worked on a "high" (maybe we could say >1B requests per day as a starting point) scale distributed service and think it's trivial to run one with more than 2 nines. You see comments all the time here mentioning that their own desktop at home is achieving more than that which belies deep misunderstanding of how systems are measured. Or t…

Well ackchually.. I get that large scale systems pose their own challenges on their own, but it also matters what's the smallest isolable unit.

What I mean by this is a CDN consists of nodes that are horizontally replicable and don't really talk to each other, and thus are easy to run even at scale.

In contrast, something like a bank or social media isn't really reducible - every user needs to be able to interact with every other user in a consistent manner.

So running a midsize bank's backend which processes 10m transactions per day, might be as if not more complex (all consistent, repeatable, and must never fail), that having a product which is a 10-10k org's IT infra replicated a thousand times.

And yes, lots of people have worked at banks and other fintech companies of this scale, including me.

I am not an expert, as I never worked on the 'core' systems but I know folks who did, and everyone told me there's an arcane database monolith that sits at the heart of these, very expensive and exotic big box SW & HW (at least for us unwashed rubes used to EC2 instances)

Re: Salesforce Global Outage

#156
post #66

Despite all of the snark here, in my experience Salesforce SRE team is quite competent. The engineering challenges of running a large PaaS - not just with own apps, but with millions of customer-written apps running on it - are quite interesting, and sadly things happen. The status page makes sense to actual customers, it's the particular "pods" where a given service runs.

As a previous SRE at Cloudflare, I'll never shit-talk fellow SREs at big companies.

The level of scale and complexity a big tech SRE has to deal with on a constant day-to-day is a very imbalanced proposition. A lot of people, in my experience, are not fully comprehending.

You have to be a jack-of-all-trades and a master of all.

Re: Salesforce Global Outage

#157
post #66

Despite all of the snark here, in my experience Salesforce SRE team is quite competent. The engineering challenges of running a large PaaS - not just with own apps, but with millions of customer-written apps running on it - are quite interesting, and sadly things happen. The status page makes sense to actual customers, it's the particular "pods" where a given service runs.

What I'm curious about is why it is a single-PaaS; I'd have expected Salesforce to have the customers quite isolated so the chance of bringing down multiple customers at once was much smaller.

You still have fleet wide management which can cause issues. Plus there are always a few core services like queues, authz, presentation layer.

Also, mono-tenant architectures is no golden bullet either. Such architecture (often coming from a formerly on-premise product that was SaaS-ified) can easily become hell to operate as it multiplies the integration points (DB parameters, URLs, allowlists, etc).

It's also quite wasteful in terms of resource utilization and hosting costs.

Re: Salesforce Global Outage

#158

You better bet someone started their agents with a prompt "Make a salesforce clone but with 100% uptime"

"Make a salesforce clone but with 100% uptime" ⎿ You've hit your session limit · resets 2:53am (48°52.6′S, 123°23.6′W Etc/GMT+8) /upgrade to increase your usage limit.

Point Nemo :-D https://en.wikipedia.org/wiki/Pole_of_inaccessibility

Re: Salesforce Global Outage

#159

Earlier quoted context omitted.

In my entire adult lifetime (mid 40s), my ISP has never been out for five days. Compare to Github, Microsoft, Salesforce, and AWS outages that are always occurring in some fashion. Reddit is down constantly in various ways and still continues to operate as a business, public no less, so I disagree about the need to chase five nines and broadly speaking, large distributed systems that are potentially unnecessary for t…

Consider yourself lucky that you’ve never been the victim of a fiber cut. But what about if the power to your house goes out? Or what if your server blows the power supply? My entire point is that you have no redundancy in your system and you also aren’t big enough to have any pull with the vendors who can fix these types of outages so you’re basically at the mercy of your providers with no recourse. That’s why these…

> That’s why these systems are built the way they are.

Built how? Because I can state with confidence that I have cleaned up a ton of failed upgrades/zombie terraform deploys of these serverless kubernetes wonders that followed every best practice under the sun, and these things are not really considered even moderately reliable (as designed by imperfect mortals under real world conditions), meanwhile professionally, people who stand to lose a lot of money should their systems go down generally operate systems whose architectures were designed decades ago, are generally horizontally scaled monoliths, and are extremely conservative in software choice.

Also downtime often is no biggie, as long as it's planned and or don't lose (too much) customer critical data.

Like nodobody cares if your test db cluster goes down for the weekend. We even shut down our db instances to save money.

Re: Salesforce Global Outage

#160
post #52

Cause: Legacy Salesforce login service got into a resource-exhaustion cascade. Fix: Rolling some unspecified fix they proved in testing out over the fleet seemingly very slowly (After their earlier attempts to roll something out faster failed). Details at https://status.salesforce.com/incidents/20004433

I wonder if “legacy login” is the shared login gateway. It’s optional but everyone uses it. And it was flaky for an hour or so, like two months ago.

And while I haven’t actually done my homework, the push to email-based login seems like it’s really pushing people to start using their own custom domains for logging in
Post reply on HN