Live data from Hacker News

The classifiers Anthropic puts in front of Fable are too zealous

combine-lab.github.io

171–180 of 203 posts

Re: The classifiers Anthropic puts in front of Fable are too zealous

#171
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

And biology is by far the classifier's least favorite topic. It's not even close. I've had it downgrade to Opus for the following questions: "How confident are we that English and American Eels both spawn in the Sargasso Sea?" "Come up with five Zoology questions of increasing difficulty for a trivia game." "What's your favorite sarcopterygian?" My wife has some zoology-related preferences in her user instructions, a…

> "What's your favorite sarcopterygian?"

Am I reading your post correctly, this question is the prompt given to an LLM? What is anyone expecting by asking an LLM what its favorite anything is? This is a conversational prompt, so accuracy and rigor is barely applicable or expected, so downgrading to a lesser model should be acceptable. If you really want to attribute preference to an LLM, consider the downgrade to be a "this conversation is beneath my advanced n-billion parameter training".

Re: The classifiers Anthropic puts in front of Fable are too zealous

#172
post #46

Earlier quoted context omitted.

And biology is by far the classifier's least favorite topic. It's not even close. I've had it downgrade to Opus for the following questions: "How confident are we that English and American Eels both spawn in the Sargasso Sea?" "Come up with five Zoology questions of increasing difficulty for a trivia game." "What's your favorite sarcopterygian?" My wife has some zoology-related preferences in her user instructions, a…

It feels like the longtermist believers got involved in this (those are the people obsessed with garage-engineered designer viruses who have a very tenuous grasp on how biology research actually works).

No "research" is needed to produce pathogens. Catastrophic genomes are already public. All someone has to do is synthesize them, which is, in actual fact, becoming more and more trivial by the day.

The inconvenience of possible mitigation strategies has no bearing on the existence of the risk itself.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#173

Earlier quoted context omitted.

I'm working on cryptography, all from academic research papers. Started well, but it eventually got some word into its context that is on the banlist. I found that if you tell it to fire off clean Fable subagents and you instruct it to check the Claude Code billing data to check for downgrades, you can get most high-sensitivity spec/review tasks done with Fable. Most. I figure that once GPT 5.6 comes out, Anthropic w…

I have been using GLM because of this reason. I think whoever makes model ignore stupid safety thing is going to win in long run.

> safety is so stupid and made-up

> makes exact argument for why people should be super concerned about AI safety

Re: The classifiers Anthropic puts in front of Fable are too zealous

#174
post #13

Earlier quoted context omitted.

Anyone test the "Gay" jailbreak to see if it works on Fable?

That wasn't even effective on ChatGPT because the results were not detailed enough, at least with Meth, in my very short testing and based on the examples.

[deleted]

Re: The classifiers Anthropic puts in front of Fable are too zealous

#175

To summarize: the classifiers Anthropic puts in front of Fable are way, way too zealous and have way too many false positives. From my experience, the model itself is very useful when it isn't refusing any of your prompts.

How do you know? Is it really that obvious of a tell between Opus and Fable? My understanding is that they silently downgrade you.

No, there's a banner that says you've been downgraded.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#176

Earlier quoted context omitted.

And biology is by far the classifier's least favorite topic. It's not even close. I've had it downgrade to Opus for the following questions: "How confident are we that English and American Eels both spawn in the Sargasso Sea?" "Come up with five Zoology questions of increasing difficulty for a trivia game." "What's your favorite sarcopterygian?" My wife has some zoology-related preferences in her user instructions, a…

> "What's your favorite sarcopterygian?" Am I reading your post correctly, this question is the prompt given to an LLM? What is anyone expecting by asking an LLM what its favorite anything is? This is a conversational prompt, so accuracy and rigor is barely applicable or expected, so downgrading to a lesser model should be acceptable. If you really want to attribute preference to an LLM, consider the downgrade to be…

I think the intent was just to show how sensitive the classifier is. If it flags prompts that simple, there's no hope for anything biology related at all really.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#178
post #56
post #47

Earlier quoted context omitted.

Anthropic's TOS clearly says they don't want to facilitate any sort of distillation, it's not a stretch to think they will limit any sort of learning on improving other models.

Literally pulling the ladder up. Disgusting behavior. I like the product, I hate the company. I can't wait for competition.

Competition is here. I personally prefer Codex. Opencode with a variety of Chinese models is also just fine for 80% of my use cases.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#179

I've found in my current work on a security auditing harness and benchmarks, both Fable and Opus are useless. I recently switched to using GPT for Nelson and the security benchmarks I've been doing because Opus started refusing to do the work. I guess I probably could also use GLM or DeepSeek or MiMo, and I'll probably do some experiments to see the shape of all of their guardrails in this area soon, now that I see i…

This underscores a huge risk of broad agentic adoption in an enterprise. Your engineers atrophy and if the agent provider decides to squeeze you, you’re SOL.
Post reply on HN