This honestly just reads as “this model failed exactly where the company said it would but I’m very special and deserve special treatment rather than the same overactive guardrails I and everyone else were told we would get.”
The classifiers Anthropic puts in front of Fable are too zealous
31–40 of 203 posts
Re: The classifiers Anthropic puts in front of Fable are too zealous
#32Re: The classifiers Anthropic puts in front of Fable are too zealous
#33This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…
I've had it downgrade to Opus for the following questions:
"How confident are we that English and American Eels both spawn in the Sargasso Sea?"
"Come up with five Zoology questions of increasing difficulty for a trivia game."
"What's your favorite sarcopterygian?"
My wife has some zoology-related preferences in her user instructions, and she had it downgrade to Opus after prompting it with: "plant."
Re: The classifiers Anthropic puts in front of Fable are too zealous
#34Do we think that someone at Anthropic, OpenAI, the government... has access to SOTA models without censorship? "How do I build an effective weapon?", "How do I effectively control the masses?"... It's very concerning that we get the nerfed models but you know that somewhere, people with a lot of resources have access to the raw, uncensored, probably more powerful models. The sprint toward AGI looks even more dangerou…
I think the problem here is that LLMs aren’t really “intelligence models” but more like “knowledge models”. LLMs don’t “think”, they just use a clever trick to make it seem like they do. I might not understand a lot about current state of AI, but that’s what they seem to be. Give it information and ask to organise it and make links, and they’ll do it, but that’s it, they don’t continually try to get out of the knowle…
I also never understand what the difference between a thinking trick, and "real" thinking is supposed to be.
Re: The classifiers Anthropic puts in front of Fable are too zealous
#35Re: The classifiers Anthropic puts in front of Fable are too zealous
#36Re: The classifiers Anthropic puts in front of Fable are too zealous
#37I'm a bioinformatician
Re: The classifiers Anthropic puts in front of Fable are too zealous
#38So you don't know what it should do, you may not even know what you would do, you don't necessarily know what's happening, and can't predict what will happen. How do you align that?
Seems like these overly sensitive filters are responding to this difficulty.
Re: The classifiers Anthropic puts in front of Fable are too zealous
#39Terrible title. Should be "Fable's guard rails are way too sensitive", which I don't think you can really blame Anthropic for. They likely had to whack them way up so it would block whatever trivial stuff got demoed to the government. I would expect them to dial down the sensitivity in a few months when nobody is looking.
I don't think it's as much when no one is looking, but instead when the broad industry SOTA, particularly Chinese models that the US government has zero control over, has advanced enough that it's security theatre restricting it.
Re: The classifiers Anthropic puts in front of Fable are too zealous
#40This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…
If your prompt has to do with those areas, yes. I haven't seen a single refusal yet. Reportedly the biology guiderails are particularly strict.