Live data from Hacker News

Salesforce Global Outage

status.salesforce.com

81–90 of 184 posts

Re: Salesforce Global Outage

#81
post #54

Have you tried turning it off and then on again? > We're no longer pursuing restarts as a path to remediation. Oh you have

Kind of surprised they admit they're going to try restarting and see what happens. I'm sure it happens everywhere but nobody admits it. > We've attempted a rolling restart on one of the impacted instances to see if that resolves the issue. At least it didn't fix the problem so they can actually start finding the real cause. > We're no longer pursuing restarts as a path to remediation. Why isn't the AI they sell telli…

I don't know, that reads exactly like an AI troubleshooter working through a plan without the implicit contextual understanding an experienced human might bring to either the actions or the communications.

"Oops, we forgot to tell it that this is the hyperscaled Salesforce production environment and that its choices need to project competence and consider brand embarrassment. WILLFIX"

Re: Salesforce Global Outage

#83
post #54

Earlier quoted context omitted.

Kind of surprised they admit they're going to try restarting and see what happens. I'm sure it happens everywhere but nobody admits it. > We've attempted a rolling restart on one of the impacted instances to see if that resolves the issue. At least it didn't fix the problem so they can actually start finding the real cause. > We're no longer pursuing restarts as a path to remediation. Why isn't the AI they sell telli…

"Yeah, I'm with Rob. Just let's reboot and see what happens"

"If that doesn't work, clear the cache and reboot again."

Re: Salesforce Global Outage

#85
post #51

Unplanned outage timing is never good but this is really not good. https://www.salesforce.com/dreamforce/ Sept 15-17

This is what happens when more than half the company is away attending the Salesforce cult-indoctrination stuff while spending all their bandwidth making customers/partners feel good.... The stuff that matters to keep the lights on gets overlooked.

I don’t think engineering and SRE of the organizing company are ever invited to those events. They’re mainly for marketing and sales (which includes solution architects).

Re: Salesforce Global Outage

#86
post #52

Cause: Legacy Salesforce login service got into a resource-exhaustion cascade. Fix: Rolling some unspecified fix they proved in testing out over the fleet seemingly very slowly (After their earlier attempts to roll something out faster failed). Details at https://status.salesforce.com/incidents/20004433

I wonder if “legacy login” is the shared login gateway.

It’s optional but everyone uses it. And it was flaky for an hour or so, like two months ago.

Re: Salesforce Global Outage

#87
post #73
post #66

Despite all of the snark here, in my experience Salesforce SRE team is quite competent. The engineering challenges of running a large PaaS - not just with own apps, but with millions of customer-written apps running on it - are quite interesting, and sadly things happen. The status page makes sense to actual customers, it's the particular "pods" where a given service runs.

I honestly don’t get the snark. The status page has: Seemingly meaningful IDs Search Region filter Email update signup Predictable URLs for instance status so they can be deep linked in runbooks What appears to be the actual live instance status . What appears to be the actual live service status in each instance. An update log with frequent detailed updates.

[flagged]

Re: Salesforce Global Outage

#88
post #51

Unplanned outage timing is never good but this is really not good. https://www.salesforce.com/dreamforce/ Sept 15-17

This is what happens when more than half the company is away attending the Salesforce cult-indoctrination stuff while spending all their bandwidth making customers/partners feel good.... The stuff that matters to keep the lights on gets overlooked.

this event is only for customers. its not a company event.

Re: Salesforce Global Outage

#89

Kind of ironic. Salesforce is basically one of the major spiritual grandfathers of Slop. It is not uncommon in production systems to find that objects like Contact and Account have hundreds of custom fields. Sometimes, you find out that several of them have the same meaning and semantics, but were used at different times. Digging out you discover that some Marketing guy that used to work at the company did some task…

Something tells me Troy the Salesforce Admin/BD Analyst did not cause the SAAS infrastructure to go down.

And I think you're confusing crud with slop.

Re: Salesforce Global Outage

#90
post #54

Have you tried turning it off and then on again? > We're no longer pursuing restarts as a path to remediation. Oh you have

Kind of surprised they admit they're going to try restarting and see what happens. I'm sure it happens everywhere but nobody admits it. > We've attempted a rolling restart on one of the impacted instances to see if that resolves the issue. At least it didn't fix the problem so they can actually start finding the real cause. > We're no longer pursuing restarts as a path to remediation. Why isn't the AI they sell telli…

Restart should be a very last emergency step, as if it works, a restart often might wipe out evidence of why.

So hopefully it's not done often.

Post reply on HN