Live data from Hacker News

The classifiers Anthropic puts in front of Fable are too zealous

combine-lab.github.io

131–140 of 203 posts

Re: The classifiers Anthropic puts in front of Fable are too zealous

#131
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

I got downgraded to Opus for asking "What is a cell?" that's all, single message, instant downgrade.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#133

Earlier quoted context omitted.

And biology is by far the classifier's least favorite topic. It's not even close. I've had it downgrade to Opus for the following questions: "How confident are we that English and American Eels both spawn in the Sargasso Sea?" "Come up with five Zoology questions of increasing difficulty for a trivia game." "What's your favorite sarcopterygian?" My wife has some zoology-related preferences in her user instructions, a…

Plants can be toxic. Be grateful if Fable doesn't report your wife for terrorism. Maybe it can help you identify where your life went off track.

Is this... Is this England?

Re: The classifiers Anthropic puts in front of Fable are too zealous

#134
post #128

I am wondering about the author's allegation that there is a user filter, not just a prompt filter. Of course it could also be the case that it is just a prompt filter, but Fable sees memories from the authors' prior sessions that cause a rejection. I wonder if the author could control for this is in some way, if Claude lets you run isolated session without memory access.

Author mentions trying incognito mode without success.

I also tried their strictly mathematical problem description and got filtered 5/5 times.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#135
post #89

Earlier quoted context omitted.

> the very thing Anthropic says it's not good for Where? Certainly not in its announcement, for one: https://platform.claude.com/docs/en/about-claude/models/intr... No "don't use this for X".

I thought it was the very first line of the product announcement, where they defined what it was they were calling "Fable" as opposed to "Mythos" in the first place: https://www.anthropic.com/news/claude-fable-5-mythos-5 > Today we’re launching Claude Fable 5: a Mythos-class model that we’ve made safe for general use. It then goes on to a lengthy and detailed section outlining the safety considerations: https://www.a…

But none of which suggest that it is not useful for math or theoretical CS tasks. The biology classifier is so miscalibrated so as to render the model useless for biology; and yes, they hint at that on the label (but not the extent of it). However, there is no description or suggestion that it is so miscalibrated that it offers up refusals in completely innocuous theoretical tasks. If it is simply a state-of-the-art model for coding, and frontend design, so be it, but at least they should be honest about that.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#136
The only way for them to release Fable is with this stuff in front. Overall, the experience is fine. It dumps me down transparently to Opus if it has a problem and does whatever it can otherwise. The fact that they were banned from offering it to people means that they have to be over-safe. This is a classic behavioral adaptation so I don't blame them. I can still find utility.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#138

I'm a medical physicist. I literally haven't been able to get Fable to answer a question I have written -- all of my work is verboten. I have however asked Claude Code (opus 4.8) to ultracode "a Fable oracle that in a digraphed, clean content, isolated environment with a minimally scoped working codebase. Ask the model at the start and the end to report exactly what its version string is. If it is not claude-fable-5,…

My theory is that they included/didn't align-away extra biology/chemistry in the training in preparation to offer an unrestricted/less restricted model to pharmaceutical companies/trusted partners. This would necessarily require a filter between the, now more "dangerous", model.

I always assumed this would be the eventual way to manage high intelligent/"dangerous" models, since all evidence shows that alignment makes them stupid: leave the actual model on the "too dangerous for the public" side, and put a censor between. When I've mentioned this a few years ago, people said this would be too expensive, but I think everyone underestimated the amount of money being thrown at all of this. :)

Re: The classifiers Anthropic puts in front of Fable are too zealous

#139
post #115

So is this the end? Are we at that point in time where ordinary people are not allowed to use more advanced models? If so this happened sooner than expected. After that point only priveleged few will access and make use of more advanced AI. Public’s access will be restricted, limited and controlled. This will only add to the power asymmetry.

I do think it's the end of unlimited access, but I'm not terribly worried about power asymmetry as such, at least not unless they hit superintelligence and none of this matters. It's not as though Big Chemistry is oppressing us all because they can order industrial acids we can't. There are strong profit motives for model providers to ensure the advanced stuff is meaningfully available, and strong political motives for them not to be perceived as picking individual winners and losers.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#140
I wonder how this plays into Anthropic's legal holds:

The retention schedule behind it:

Deleted conversations: removed from your chat history immediately, but kept on back-end systems for up to 30 days before permanent deletion. Flagged inputs and outputs (Usage Policy violation): retained up to 2 years. Trust-and-safety classification scores (on flagged sessions): retained up to 7 years. API logs: 7 days by default (as of September 14, 2025), extendable to 30 days via a DPA. Zero Data Retention (qualifying enterprise): inputs and outputs aren't stored after the API response returns, though safety classifier results are still retained even here.

Post reply on HN