Earlier quoted context omitted.
Then what is it they are trying to guard against, if its not simply protecting their moat ahead of their IPO? Because from the outside, their behavior looks like a situation of "What if Microsoft/Apple put controls in place to make it impossible to develop an operating system using their OS?"
They are trying to guard against other people building ASI before they do because they think they are uniquely safety oriented relative to their competitors. Frankly, based on my knowledge of Anthropic and the people who work there, they are very possibly right. They care a ton about this in a way that is difficult for people outside this bubble to understand.
Anthropic apologizes for invisible Claude Fable guardrails
141–150 of 489 posts
Re: Anthropic apologizes for invisible Claude Fable guardrails
#142Earlier quoted context omitted.
I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.
I asked it to analyse my architecture and find any security issues and it did it perfectly, first identified the issues & then fixed them. Not sure why my prompt managed to get through the guardrails
Just before asking for approval to run, it said one thing it wanted to "flag before running" was "Rate-limit and auth testing against prod will generate some 4xx noise in Railway logs and could trip the form rate limiter — harmless, but saying it now."
Ok fine, I said go for it, and it says:
"Running it. Quick recon first (prod URLs + the prior-findings baseline), then I'll fan out the audit tracks with adversarial verification."
Immediately after, I got the Fable warning about how it can't continue because of safety concerns, switching to Opus. In the end, Opus did a good job thanks to whatever Fable suggested doing. Things were fixed that Opus missed in a security/performance audit just the week prior. But what surprised me is that it used 55 agents. Burned 80% of my 5-hour window in 15 minutes (5x Max plan). I've never had Opus do that before on these audits.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#143Earlier quoted context omitted.
I think it's also worth noting that EA is closely linked to utilitarianism. Most of the pitfalls that people see in EA are the same pitfalls that are classic to utilitarianism, a la "we're going to do this thing we know is locally-bad, because we have a lot of confidence in other effects that are universally-good".
It's important to separate objections to utilitarianism from the obvious fact that it can very be hard to correctly apply the utilitarian calculus. It's partly because of this difficulty that most classical utilitarians thought that people should generally follow commonsense morality and not try to directly apply the utilitarian calculus (which then led to the charge of paternalism and teaching one morality to the ma…
Even people who say they are deontologists often slip back into utilitarian arguments when they're not careful — for example, when arguing Kant's categorical imperative against lying, they slip into talking about the local benefits vs. overall harms.
The real gap, as you've said, is more about overconfidence in one's utilitarian calculus for distal vs. proximal moral outcomes. An average Joe is likely to give a lot more weight to the moral outcomes that rely on local information and affect his friends and family. The characterization of an "EA" — whether fair or not — is that they're much more likely to use a lot more explicit moral calculus and attempt to correct for proximal vs. distal biases.
In a way its very similar to Sowell's arguments about the informational economics of a distributed market vs. a central planner.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#144Earlier quoted context omitted.
No, it was not clear. No one expects that a tool they pay for and use professionally to purposefully sabotage their work. You’re excusing their unhinged behavior. https://xcancel.com/hammer_mt/status/2064839924398825798
Making excuses for billion+ dollar companies' behavior is one of the most common HN comment section pastimes.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#145Earlier quoted context omitted.
Nothing, they are just trying to scare monger the public and prime the pump for a massive bailout when it crashes out because apparently China are the big bad meanies.
You'd be fine if the PRC gets to ASI first? That's an interesting opinion.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#146Earlier quoted context omitted.
Yeah, I cancelled my Claude subscription yesterday after learning about their attitude of intentionally sabotaging their paying customers. Especially after trying Fable yesterday for some benign projects and being unimpressive relative to opus. Rolling it back is the right move, but I’m still not convinced that using them is in my best interest anymore, I’m investigating open source cloud providers now.
Opus is nowhere close to Fable. Fable feels at least one generation ahead to me. https://x.com/hyperagentapp/status/2064396004032463157 Edit: OpenAI will launch a similar model soon and I can't wait. We are entering a new era of agents.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#147Earlier quoted context omitted.
It's an interesting assumption. The idea behind this with nukes was that we'd like to nuke Germany before they could nuke us. Even after we defeated Germany, we nuked Japan even though they had no possibility of getting their own nukes. The nuclear 'race' was based on the premise that the winner could use it to destroy all other racers (a faulty assumption, see the USSR among others). I will charitably assume Anthrop…
Do you believe the current situation is more akin to the race to the first nukes , where no one could know for sure the other competitors were even racing... or is it more similar to the Cold War, where there were obviously competitors engaged in the race? And yes, agreed the equilibrium dynamics for AGI are very different (and far harder to predict) than nukes. That sounds like a good reason to be sure we get there…
Re: Anthropic apologizes for invisible Claude Fable guardrails
#148Earlier quoted context omitted.
Yeah, I cancelled my Claude subscription yesterday after learning about their attitude of intentionally sabotaging their paying customers. Especially after trying Fable yesterday for some benign projects and being unimpressive relative to opus. Rolling it back is the right move, but I’m still not convinced that using them is in my best interest anymore, I’m investigating open source cloud providers now.
Opus is nowhere close to Fable. Fable feels at least one generation ahead to me. https://x.com/hyperagentapp/status/2064396004032463157 Edit: OpenAI will launch a similar model soon and I can't wait. We are entering a new era of agents.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#149What's interesting is they say they'll change this to an explicit refusal in a few days, which seems too fast for them to retrain Fable/Mythos itself, so implies that this was always a filter in front of the model, and judging by how crude their "safety" filter is, this "might compete with us" filter is not going to be any better.
I also wonder who's paying for the tokens consumed by the filter (presumably also an LLM) - is that now factored into the input tokens cost? Hopefully(?) it is an LLM not just a regex like Claude Code's "sentiment" (swear) detector.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#150Earlier quoted context omitted.
You'd be fine if the PRC gets to ASI first? That's an interesting opinion.
PRC labs reportedly aren't even thinking about getting to ASI, much less trying. They think of AI as a technology that can provide utility across the board even without anything like superhuman smarts.
It smells of paranoia.