Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

141–150 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#141

Earlier quoted context omitted.

Then what is it they are trying to guard against, if its not simply protecting their moat ahead of their IPO? Because from the outside, their behavior looks like a situation of "What if Microsoft/Apple put controls in place to make it impossible to develop an operating system using their OS?"

They are trying to guard against other people building ASI before they do because they think they are uniquely safety oriented relative to their competitors. Frankly, based on my knowledge of Anthropic and the people who work there, they are very possibly right. They care a ton about this in a way that is difficult for people outside this bubble to understand.

Define safety oriented.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#142
post #15

Earlier quoted context omitted.

I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.

I asked it to analyse my architecture and find any security issues and it did it perfectly, first identified the issues & then fixed them. Not sure why my prompt managed to get through the guardrails

I asked Fable to plan a security & performance audit of my website. It said it would check SSR & origin attack surface, CMS content injection, Strapi API surface, etc.

Just before asking for approval to run, it said one thing it wanted to "flag before running" was "Rate-limit and auth testing against prod will generate some 4xx noise in Railway logs and could trip the form rate limiter — harmless, but saying it now."

Ok fine, I said go for it, and it says:

"Running it. Quick recon first (prod URLs + the prior-findings baseline), then I'll fan out the audit tracks with adversarial verification."

Immediately after, I got the Fable warning about how it can't continue because of safety concerns, switching to Opus. In the end, Opus did a good job thanks to whatever Fable suggested doing. Things were fixed that Opus missed in a security/performance audit just the week prior. But what surprised me is that it used 55 agents. Burned 80% of my 5-hour window in 15 minutes (5x Max plan). I've never had Opus do that before on these audits.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#143

Earlier quoted context omitted.

I think it's also worth noting that EA is closely linked to utilitarianism. Most of the pitfalls that people see in EA are the same pitfalls that are classic to utilitarianism, a la "we're going to do this thing we know is locally-bad, because we have a lot of confidence in other effects that are universally-good".

It's important to separate objections to utilitarianism from the obvious fact that it can very be hard to correctly apply the utilitarian calculus. It's partly because of this difficulty that most classical utilitarians thought that people should generally follow commonsense morality and not try to directly apply the utilitarian calculus (which then led to the charge of paternalism and teaching one morality to the ma…

Yes that's a very good point.

Even people who say they are deontologists often slip back into utilitarian arguments when they're not careful — for example, when arguing Kant's categorical imperative against lying, they slip into talking about the local benefits vs. overall harms.

The real gap, as you've said, is more about overconfidence in one's utilitarian calculus for distal vs. proximal moral outcomes. An average Joe is likely to give a lot more weight to the moral outcomes that rely on local information and affect his friends and family. The characterization of an "EA" — whether fair or not — is that they're much more likely to use a lot more explicit moral calculus and attempt to correct for proximal vs. distal biases.

In a way its very similar to Sowell's arguments about the informational economics of a distributed market vs. a central planner.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#144

Earlier quoted context omitted.

No, it was not clear. No one expects that a tool they pay for and use professionally to purposefully sabotage their work. You’re excusing their unhinged behavior. https://xcancel.com/hammer_mt/status/2064839924398825798

Making excuses for billion+ dollar companies' behavior is one of the most common HN comment section pastimes.

Only second to making intellectually dishonest criticisms of perceived behaviours

Re: Anthropic apologizes for invisible Claude Fable guardrails

#145

Earlier quoted context omitted.

Nothing, they are just trying to scare monger the public and prime the pump for a massive bailout when it crashes out because apparently China are the big bad meanies.

You'd be fine if the PRC gets to ASI first? That's an interesting opinion.

PRC labs reportedly aren't even thinking about getting to ASI, much less trying. They think of AI as a technology that can provide utility across the board even without anything like superhuman smarts.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#146

Earlier quoted context omitted.

Yeah, I cancelled my Claude subscription yesterday after learning about their attitude of intentionally sabotaging their paying customers. Especially after trying Fable yesterday for some benign projects and being unimpressive relative to opus. Rolling it back is the right move, but I’m still not convinced that using them is in my best interest anymore, I’m investigating open source cloud providers now.

Opus is nowhere close to Fable. Fable feels at least one generation ahead to me. https://x.com/hyperagentapp/status/2064396004032463157 Edit: OpenAI will launch a similar model soon and I can't wait. We are entering a new era of agents.

ad

Re: Anthropic apologizes for invisible Claude Fable guardrails

#147

Earlier quoted context omitted.

It's an interesting assumption. The idea behind this with nukes was that we'd like to nuke Germany before they could nuke us. Even after we defeated Germany, we nuked Japan even though they had no possibility of getting their own nukes. The nuclear 'race' was based on the premise that the winner could use it to destroy all other racers (a faulty assumption, see the USSR among others). I will charitably assume Anthrop…

Do you believe the current situation is more akin to the race to the first nukes , where no one could know for sure the other competitors were even racing... or is it more similar to the Cold War, where there were obviously competitors engaged in the race? And yes, agreed the equilibrium dynamics for AGI are very different (and far harder to predict) than nukes. That sounds like a good reason to be sure we get there…

I can't really say I see a similarity to either the Manhattan Project or the Cold War. I don't see how one could apply either massive retaliation or MAD. These are private companies, they are not vested with the necessary authority to destroy anything. Even if they had it, they couldn't. You can't destroy China, they have 1.4B people, nukes, and a large part of the world's manufacturing. So multiple organizations want to do something first, that could be anything from nukes to railroads to lining up for communion wafers.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#148

Earlier quoted context omitted.

Yeah, I cancelled my Claude subscription yesterday after learning about their attitude of intentionally sabotaging their paying customers. Especially after trying Fable yesterday for some benign projects and being unimpressive relative to opus. Rolling it back is the right move, but I’m still not convinced that using them is in my best interest anymore, I’m investigating open source cloud providers now.

Opus is nowhere close to Fable. Fable feels at least one generation ahead to me. https://x.com/hyperagentapp/status/2064396004032463157 Edit: OpenAI will launch a similar model soon and I can't wait. We are entering a new era of agents.

Fable is very much an incremental development over Opus, and even more incremental when properly compared to its existing counterparts GPT-Pro and Gemini Deep Research.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#149
I suppose it's an improvement, but it doesn't make the model any more useful. Anthropic are now being quite explicit that they'll choose what you can and can't use their models for, and most importantly that's not limited to any safety concerns - it includes not allowing you to work on AI (and anything else Anthropic may choose to work on).

What's interesting is they say they'll change this to an explicit refusal in a few days, which seems too fast for them to retrain Fable/Mythos itself, so implies that this was always a filter in front of the model, and judging by how crude their "safety" filter is, this "might compete with us" filter is not going to be any better.

I also wonder who's paying for the tokens consumed by the filter (presumably also an LLM) - is that now factored into the input tokens cost? Hopefully(?) it is an LLM not just a regex like Claude Code's "sentiment" (swear) detector.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#150

Earlier quoted context omitted.

You'd be fine if the PRC gets to ASI first? That's an interesting opinion.

PRC labs reportedly aren't even thinking about getting to ASI, much less trying. They think of AI as a technology that can provide utility across the board even without anything like superhuman smarts.

A lot of this lust for ASI is driven by America attempting to cling onto the power it has wielded over the world over the past 50 odd yrs.

It smells of paranoia.

Post reply on HN