Live data from Hacker News

The classifiers Anthropic puts in front of Fable are too zealous

combine-lab.github.io

41–50 of 203 posts

Re: The classifiers Anthropic puts in front of Fable are too zealous

#41
post #6

Do we think that someone at Anthropic, OpenAI, the government... has access to SOTA models without censorship? "How do I build an effective weapon?", "How do I effectively control the masses?"... It's very concerning that we get the nerfed models but you know that somewhere, people with a lot of resources have access to the raw, uncensored, probably more powerful models. The sprint toward AGI looks even more dangerou…

[deleted]

Re: The classifiers Anthropic puts in front of Fable are too zealous

#44
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

If your prompt has to do with those areas, yes. I haven't seen a single refusal yet. Reportedly the biology guiderails are particularly strict.

I get it every time I touch the login path

Re: The classifiers Anthropic puts in front of Fable are too zealous

#46
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

And biology is by far the classifier's least favorite topic. It's not even close. I've had it downgrade to Opus for the following questions: "How confident are we that English and American Eels both spawn in the Sargasso Sea?" "Come up with five Zoology questions of increasing difficulty for a trivia game." "What's your favorite sarcopterygian?" My wife has some zoology-related preferences in her user instructions, a…

It feels like the longtermist believers got involved in this (those are the people obsessed with garage-engineered designer viruses who have a very tenuous grasp on how biology research actually works).

Re: The classifiers Anthropic puts in front of Fable are too zealous

#47

Fable was refusing to patch vllm for me when trying to get mtp to work on r9700 gpus. Kept on bumping down to opus. Tried to really sanitize my prompts and everything but it seemed intrinsically prohibited from doing this sort of work. I guess it’s useful for making inane one shot games and websites, lol.

Anthropic's TOS clearly says they don't want to facilitate any sort of distillation, it's not a stretch to think they will limit any sort of learning on improving other models.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#48

Fable was refusing to patch vllm for me when trying to get mtp to work on r9700 gpus. Kept on bumping down to opus. Tried to really sanitize my prompts and everything but it seemed intrinsically prohibited from doing this sort of work. I guess it’s useful for making inane one shot games and websites, lol.

That's where the free marketing comes from.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#50

I'm curious what the state of alignment research is. My gut says this is basically impossible. People have different moral frameworks. Each individual probably has an inconsistent moral framework. Even granting perfect consistency, applying these typically requires some knowledge of reality. And these LLM / harness combos are turing complete. So you don't know what it should do, you may not even know what you would d…

it's anthropics moral framework that matters, not the myriad of moral frameworks of the individual users
Post reply on HN