Live data from Hacker News

The classifiers Anthropic puts in front of Fable are too zealous

combine-lab.github.io

1–10 of 203 posts

Re: The classifiers Anthropic puts in front of Fable are too zealous

#3
Terrible title. Should be "Fable's guard rails are way too sensitive", which I don't think you can really blame Anthropic for. They likely had to whack them way up so it would block whatever trivial stuff got demoed to the government.

I would expect them to dial down the sensitivity in a few months when nobody is looking.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#5
This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness.

e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, despite only being very marginally, tangentially, somewhat related to biology.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#6
Do we think that someone at Anthropic, OpenAI, the government... has access to SOTA models without censorship? "How do I build an effective weapon?", "How do I effectively control the masses?"...

It's very concerning that we get the nerfed models but you know that somewhere, people with a lot of resources have access to the raw, uncensored, probably more powerful models. The sprint toward AGI looks even more dangerous when you think about who will be gaining access to it first. I do believe the goal is to pull away from the rest of humanity in a near trans-humanistic state. Are we ready for that and how do we counter it?

Re: The classifiers Anthropic puts in front of Fable are too zealous

#7
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

If your prompt has to do with those areas, yes. I haven't seen a single refusal yet.

Reportedly the biology guiderails are particularly strict.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#8
I have only really used Fable as a final pass on something. A "Take a look at everything we did so far, and make sure we didn't forget something" kind of review prompt.

But it is a huge waste of money for most coding tasks. Opus is still overkill most of the time, too.

Post reply on HN