Live data from Hacker News

The classifiers Anthropic puts in front of Fable are too zealous

combine-lab.github.io

31–40 of 203 posts

Re: The classifiers Anthropic puts in front of Fable are too zealous

#31

This honestly just reads as “this model failed exactly where the company said it would but I’m very special and deserve special treatment rather than the same overactive guardrails I and everyone else were told we would get.”

I just think that Anthropic's usage of the word "classifier", which implies a minimum level of intelligence, was very misleading. Fact is, you cannot use Fable for anything remotely connected to even elementary school biology or medical topics. There is no attempt whatsoever to distinguish between legitimate and dangerous tasks, except an extremely broad and non-specific rejection of anything related to security or biology.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#33
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

And biology is by far the classifier's least favorite topic. It's not even close.

I've had it downgrade to Opus for the following questions:

"How confident are we that English and American Eels both spawn in the Sargasso Sea?"

"Come up with five Zoology questions of increasing difficulty for a trivia game."

"What's your favorite sarcopterygian?"

My wife has some zoology-related preferences in her user instructions, and she had it downgrade to Opus after prompting it with: "plant."

Re: The classifiers Anthropic puts in front of Fable are too zealous

#34
post #6

Do we think that someone at Anthropic, OpenAI, the government... has access to SOTA models without censorship? "How do I build an effective weapon?", "How do I effectively control the masses?"... It's very concerning that we get the nerfed models but you know that somewhere, people with a lot of resources have access to the raw, uncensored, probably more powerful models. The sprint toward AGI looks even more dangerou…

I think the problem here is that LLMs aren’t really “intelligence models” but more like “knowledge models”. LLMs don’t “think”, they just use a clever trick to make it seem like they do. I might not understand a lot about current state of AI, but that’s what they seem to be. Give it information and ask to organise it and make links, and they’ll do it, but that’s it, they don’t continually try to get out of the knowle…

When you watch it solve complex problems and use the browser and do internet searches, and use the entire surface area of the console tools on a linux box every day the idea that there are no major Homomorphisms with biological thinking is just completely out of the question.

I also never understand what the difference between a thinking trick, and "real" thinking is supposed to be.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#38
I'm curious what the state of alignment research is. My gut says this is basically impossible. People have different moral frameworks. Each individual probably has an inconsistent moral framework. Even granting perfect consistency, applying these typically requires some knowledge of reality. And these LLM / harness combos are turing complete.

So you don't know what it should do, you may not even know what you would do, you don't necessarily know what's happening, and can't predict what will happen. How do you align that?

Seems like these overly sensitive filters are responding to this difficulty.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#39
post #3

Terrible title. Should be "Fable's guard rails are way too sensitive", which I don't think you can really blame Anthropic for. They likely had to whack them way up so it would block whatever trivial stuff got demoed to the government. I would expect them to dial down the sensitivity in a few months when nobody is looking.

> I would expect them to dial down the sensitivity in a few months when nobody is looking.

I don't think it's as much when no one is looking, but instead when the broad industry SOTA, particularly Chinese models that the US government has zero control over, has advanced enough that it's security theatre restricting it.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#40
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

If your prompt has to do with those areas, yes. I haven't seen a single refusal yet. Reportedly the biology guiderails are particularly strict.

It’s happy to work on our backend repository. It refuses to work on our infrastructure (Terraform) repo.
Post reply on HN