Live data from Hacker News

The classifiers Anthropic puts in front of Fable are too zealous

combine-lab.github.io

141–150 of 203 posts

Re: The classifiers Anthropic puts in front of Fable are too zealous

#141
I've found in my current work on a security auditing harness and benchmarks, both Fable and Opus are useless. I recently switched to using GPT for Nelson and the security benchmarks I've been doing because Opus started refusing to do the work. I guess I probably could also use GLM or DeepSeek or MiMo, and I'll probably do some experiments to see the shape of all of their guardrails in this area soon, now that I see it's more than one model that behaves this way (Gemini in Antigravity also refuses any security auditing task, even as simple as "find security bugs").

I blogged about it: https://swelljoe.com/post/why-i-had-to-switch-to-gpt/

Re: The classifiers Anthropic puts in front of Fable are too zealous

#142
post #43

[flagged]

I wanted to use Fable to discuss a philosophical topic, but halfway through I used the word "cell" and got deflected to Opus.

On the other hand Opus has this awful adversarial-teacher vibe. It pushes back for no useful reason, talks down to you, and acts like it has to prove itself by grading and correcting everything. Instead of working with your claim, it reframes it, declares what the "real" issue is, then tells you what you failed to do.

So Fable refuses me and I can't stand Opus. Nice one, Anthropic, I need to downgrade my subscription.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#143
post #81
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

I've been working with it heavily since its first release. I use it for software architects, complex debugging and some development and I have not had it refuse or downgrade even once.

That’s interesting. I find it completely unusable for even simple reviews of existing project documentation that I wrote for an iOS app that isn’t even in public distribution.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#144
post #43

[flagged]

I havent used fable, but does it cost you the same when it downgrades to opus?

I believe they confirmed, on twitter or somewhere I frustratingly can't find, that downgrades are charged at the correct opus rate, after a user asked and was told "either that's how it works or it's a bug"

Re: The classifiers Anthropic puts in front of Fable are too zealous

#145

Earlier quoted context omitted.

When you watch it solve complex problems and use the browser and do internet searches, and use the entire surface area of the console tools on a linux box every day the idea that there are no major Homomorphisms with biological thinking is just completely out of the question. I also never understand what the difference between a thinking trick, and "real" thinking is supposed to be.

I used to agree with you but overtime I’ve changed my mind. For reference I created predictive linguistics at Google in the first products and this is a many order scale up of that, with new complexities of course. The best analogy I can give you is that it is a really advanced synthesis machine, which looks like human thought but is more of a hyper advance “replay” of human thought in various contexts. Where you beg…

I feel like that's more saying they can't train on the fly, and also that serializing spatial data and world models is something we haven't really done fully.

For me all neural networks synthetic or otherwise are replay machines or stream prediction machines. Nerve signals in, and nerve signals out. If I create output signals to the muscle nerves like this when my eyes see signals like that, good things happen, i get a reward, so it happens again the next time. We have a a more complex messier architecture, but it seems pretty much the same in the input and outputs being linear signals.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#146
post #126
post #69

Earlier quoted context omitted.

My experience, too. I work on nothing in any way related to cybersecurity or biology. I asked it a few purely mathematical questions, it refused immediately. Before the export embargo I did get it to look at some hairy problems and the output was genuinely useful...

Were your question in math areas related to ML? They also restrict model development and research pretty heavily.

No. Robust control theory was one case. Dynamical systems another.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#147
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

And biology is by far the classifier's least favorite topic. It's not even close. I've had it downgrade to Opus for the following questions: "How confident are we that English and American Eels both spawn in the Sargasso Sea?" "Come up with five Zoology questions of increasing difficulty for a trivia game." "What's your favorite sarcopterygian?" My wife has some zoology-related preferences in her user instructions, a…

I had a file that had a couple places where vars were named DNA and got just total refusals during the first launch. Came away thinking the model was total trash. The guardrail classifiers are for sure total trash.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#148

To summarize: the classifiers Anthropic puts in front of Fable are way, way too zealous and have way too many false positives. From my experience, the model itself is very useful when it isn't refusing any of your prompts.

How do you know? Is it really that obvious of a tell between Opus and Fable? My understanding is that they silently downgrade you.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#149

I wonder how this plays into Anthropic's legal holds: The retention schedule behind it: Deleted conversations: removed from your chat history immediately, but kept on back-end systems for up to 30 days before permanent deletion. Flagged inputs and outputs (Usage Policy violation): retained up to 2 years. Trust-and-safety classification scores (on flagged sessions): retained up to 7 years. API logs: 7 days by default…

update: Anthropic has not published the retention treatment of routing metadata, in particular whether a reroute counts only as caution (the 30-day safety-monitoring floor) or as a Usage Policy violation flag (the 2-year content and 7-year score horizons in Part I). That distinction is legally consequential, because a flagged Fable session could persist far longer than 30 days. The classifier's internal decision logic is also deliberately undisclosed.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#150
They're not just "too zealous", they're ludicrous.

I've had it reject looking at pages served from my local network because it "can't find it with my search tool" and had "ethical concerns about consent for access".

The People's AI Concern Front has gotten the classifier they want, and it's made Claude hilariously useless. I am waiting with bated breath for their next set of revenue numbers. (And happily hand my money to competitors instead)

Post reply on HN