This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…
The classifiers Anthropic puts in front of Fable are too zealous
131–140 of 203 posts
Re: The classifiers Anthropic puts in front of Fable are too zealous
#132Re: The classifiers Anthropic puts in front of Fable are too zealous
#133Earlier quoted context omitted.
And biology is by far the classifier's least favorite topic. It's not even close. I've had it downgrade to Opus for the following questions: "How confident are we that English and American Eels both spawn in the Sargasso Sea?" "Come up with five Zoology questions of increasing difficulty for a trivia game." "What's your favorite sarcopterygian?" My wife has some zoology-related preferences in her user instructions, a…
Plants can be toxic. Be grateful if Fable doesn't report your wife for terrorism. Maybe it can help you identify where your life went off track.
Re: The classifiers Anthropic puts in front of Fable are too zealous
#134I am wondering about the author's allegation that there is a user filter, not just a prompt filter. Of course it could also be the case that it is just a prompt filter, but Fable sees memories from the authors' prior sessions that cause a rejection. I wonder if the author could control for this is in some way, if Claude lets you run isolated session without memory access.
I also tried their strictly mathematical problem description and got filtered 5/5 times.
Re: The classifiers Anthropic puts in front of Fable are too zealous
#135Earlier quoted context omitted.
> the very thing Anthropic says it's not good for Where? Certainly not in its announcement, for one: https://platform.claude.com/docs/en/about-claude/models/intr... No "don't use this for X".
I thought it was the very first line of the product announcement, where they defined what it was they were calling "Fable" as opposed to "Mythos" in the first place: https://www.anthropic.com/news/claude-fable-5-mythos-5 > Today we’re launching Claude Fable 5: a Mythos-class model that we’ve made safe for general use. It then goes on to a lengthy and detailed section outlining the safety considerations: https://www.a…
Re: The classifiers Anthropic puts in front of Fable are too zealous
#136Re: The classifiers Anthropic puts in front of Fable are too zealous
#137[flagged]
Re: The classifiers Anthropic puts in front of Fable are too zealous
#138I'm a medical physicist. I literally haven't been able to get Fable to answer a question I have written -- all of my work is verboten. I have however asked Claude Code (opus 4.8) to ultracode "a Fable oracle that in a digraphed, clean content, isolated environment with a minimally scoped working codebase. Ask the model at the start and the end to report exactly what its version string is. If it is not claude-fable-5,…
I always assumed this would be the eventual way to manage high intelligent/"dangerous" models, since all evidence shows that alignment makes them stupid: leave the actual model on the "too dangerous for the public" side, and put a censor between. When I've mentioned this a few years ago, people said this would be too expensive, but I think everyone underestimated the amount of money being thrown at all of this. :)
Re: The classifiers Anthropic puts in front of Fable are too zealous
#139So is this the end? Are we at that point in time where ordinary people are not allowed to use more advanced models? If so this happened sooner than expected. After that point only priveleged few will access and make use of more advanced AI. Public’s access will be restricted, limited and controlled. This will only add to the power asymmetry.
Re: The classifiers Anthropic puts in front of Fable are too zealous
#140The retention schedule behind it:
Deleted conversations: removed from your chat history immediately, but kept on back-end systems for up to 30 days before permanent deletion. Flagged inputs and outputs (Usage Policy violation): retained up to 2 years. Trust-and-safety classification scores (on flagged sessions): retained up to 7 years. API logs: 7 days by default (as of September 14, 2025), extendable to 30 days via a DPA. Zero Data Retention (qualifying enterprise): inputs and outputs aren't stored after the API response returns, though safety classifier results are still retained even here.