Live data from Hacker News

The classifiers Anthropic puts in front of Fable are too zealous

combine-lab.github.io

61–70 of 203 posts

Re: The classifiers Anthropic puts in front of Fable are too zealous

#62

Fable was refusing to patch vllm for me when trying to get mtp to work on r9700 gpus. Kept on bumping down to opus. Tried to really sanitize my prompts and everything but it seemed intrinsically prohibited from doing this sort of work. I guess it’s useful for making inane one shot games and websites, lol.

I was recently using self-hosted DeepSeek V4 Flash to poke around the DSpark implementation in vLLM (well outside of my domain)

I did wonder if I was doing anything Fable would have flagged - sounds like yes.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#63
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

Additionally, I thought the threat being modeled for “biology” was stuff like bioterroism- how to make anthrax, how to distribute a toxin, etc.

I don’t feel like calculating results for a trial is really in the threat model unless we think a terrorist is out there testing the efficacy of their anthrax before using it in an attack.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#64
post #46

Earlier quoted context omitted.

And biology is by far the classifier's least favorite topic. It's not even close. I've had it downgrade to Opus for the following questions: "How confident are we that English and American Eels both spawn in the Sargasso Sea?" "Come up with five Zoology questions of increasing difficulty for a trivia game." "What's your favorite sarcopterygian?" My wife has some zoology-related preferences in her user instructions, a…

It feels like the longtermist believers got involved in this (those are the people obsessed with garage-engineered designer viruses who have a very tenuous grasp on how biology research actually works).

Yeah i'm wondering how much of a role that plays in this as well.

On the one hand I could believe it's something more benign, or the usual misunderstood fear mongering making it to some political level (well make sure those users can't get online anonymously! being our current craze).

That said, chemistry and to some level physics have been the major domain of limited knowledge (chemistry because the average person could cause some damage, physics is more of a nation state issue generally).

However I do wonder if there's some legit data on "oh uh...looks like this thing you can make with easy to get and hard to regulate tools is dangerous" in the bio field. I know about the lab rats who want to just screw around in the garage, and it seems like that should be easy to hit at a supply level (much like how certain chemical compounds are just not available for civilians), but maybe there's something legit to limiting the data.

Not that this is a remotely good implementation of that. The hamfisted method does reek of some politician/bureaucrat just saying "No it can't ever return bio questions because RAR!" situation.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#65
post #27
post #6

Do we think that someone at Anthropic, OpenAI, the government... has access to SOTA models without censorship? "How do I build an effective weapon?", "How do I effectively control the masses?"... It's very concerning that we get the nerfed models but you know that somewhere, people with a lot of resources have access to the raw, uncensored, probably more powerful models. The sprint toward AGI looks even more dangerou…

Yes, of course. Until Fable even the public had practically uncensored access to SOTA anthropic models (there were classifiers - but they were very hard to hit). And I'd have to double check but I'm pretty certain the public still has uncensored access to SOTA models from google (via GCP under threat of Google ceasing to do business with you and theoretically suing you if you violate the TOS). Censorship being what t…

> It strikes me as highly unlikely that Anthropic has developed another fable-class model where the only difference is that it doesn't answer questions in that way

I'm curious why you think that's highly unlikely given the monetary incentive (or even post-monetary!) to create such a thing? I imagine there's also an arms race aspect, if you assume your enemies (whoever they are) have access to such a model, certainly those capable of creating one, would.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#66
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

And biology is by far the classifier's least favorite topic. It's not even close. I've had it downgrade to Opus for the following questions: "How confident are we that English and American Eels both spawn in the Sargasso Sea?" "Come up with five Zoology questions of increasing difficulty for a trivia game." "What's your favorite sarcopterygian?" My wife has some zoology-related preferences in her user instructions, a…

Plants can be toxic. Be grateful if Fable doesn't report your wife for terrorism. Maybe it can help you identify where your life went off track.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#68
post #8

I have only really used Fable as a final pass on something. A "Take a look at everything we did so far, and make sure we didn't forget something" kind of review prompt. But it is a huge waste of money for most coding tasks. Opus is still overkill most of the time, too.

I have used Fable to the full extend of the 20x subscription's weekly limit, for all development tasks on my iOS project.

It was working better than Opus for me. It more often implemented features well on the first try, where Opus needs a few rounds of improvements to reach a passable result.

I am not sure why it would be a waste of money "for most coding tasks", and how you could conclude so with any confidence when you did not even really use it aside from final review passes.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#69
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

My experience, too. I work on nothing in any way related to cybersecurity or biology. I asked it a few purely mathematical questions, it refused immediately.

Before the export embargo I did get it to look at some hairy problems and the output was genuinely useful...

Re: The classifiers Anthropic puts in front of Fable are too zealous

#70
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

I absolutely have been unable to use Fable for any neuroimaging work. Its fine. The other models are good enough, honestly...and while I AM annoyed that the filter is so broad, I also understand it, as I do believe that models can become dangerous as WMDs, eventually. Still, it is completely useless for me.

The only question I had was being flagged for other reasons, so I asked it a mechanical engineering question, and it was just fine with that.

Post reply on HN