Earlier quoted context omitted.
Yes, of course. Until Fable even the public had practically uncensored access to SOTA anthropic models (there were classifiers - but they were very hard to hit). And I'd have to double check but I'm pretty certain the public still has uncensored access to SOTA models from google (via GCP under threat of Google ceasing to do business with you and theoretically suing you if you violate the TOS). Censorship being what t…
> It strikes me as highly unlikely that Anthropic has developed another fable-class model where the only difference is that it doesn't answer questions in that way I'm curious why you think that's highly unlikely given the monetary incentive (or even post-monetary!) to create such a thing? I imagine there's also an arms race aspect, if you assume your enemies (whoever they are) have access to such a model, certainly…
The classifiers Anthropic puts in front of Fable are too zealous
71–80 of 203 posts
Re: The classifiers Anthropic puts in front of Fable are too zealous
#72[flagged]
there's plenty of uses for models not doing what they were made to do, but this is even worse. It's people trying to get the model to do what it was made specifically not to do!
Re: The classifiers Anthropic puts in front of Fable are too zealous
#73[flagged]
Re: The classifiers Anthropic puts in front of Fable are too zealous
#74Do we think that someone at Anthropic, OpenAI, the government... has access to SOTA models without censorship? "How do I build an effective weapon?", "How do I effectively control the masses?"... It's very concerning that we get the nerfed models but you know that somewhere, people with a lot of resources have access to the raw, uncensored, probably more powerful models. The sprint toward AGI looks even more dangerou…
I think the problem here is that LLMs aren’t really “intelligence models” but more like “knowledge models”. LLMs don’t “think”, they just use a clever trick to make it seem like they do. I might not understand a lot about current state of AI, but that’s what they seem to be. Give it information and ask to organise it and make links, and they’ll do it, but that’s it, they don’t continually try to get out of the knowle…
I can make a program that writes a stories involving Santa Claus, and I can make another program that takes the hidden script and performs certain lines... but at the end of the day I have not made him real.
Re: The classifiers Anthropic puts in front of Fable are too zealous
#75Re: The classifiers Anthropic puts in front of Fable are too zealous
#76For anyone using these models for anything remotely sensitive, keep in mind that Anthropic says [0]: > We retain inputs and outputs for up to 2 years and trust and safety classification scores for up to 7 years if your chat is flagged by our automated trust and safety systems as violating our Usage Policy. And, since those automated systems apparently have a ludicrous false-positive rate, you should assume that your…
Re: The classifiers Anthropic puts in front of Fable are too zealous
#77This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…
Additionally, I thought the threat being modeled for “biology” was stuff like bioterroism- how to make anthrax, how to distribute a toxin, etc. I don’t feel like calculating results for a trial is really in the threat model unless we think a terrorist is out there testing the efficacy of their anthrax before using it in an attack.
You can ask it elementary school grade biology trivia, or obscure facts about recently documented insect species, and both will downgrade to Opus 4.8 straight away.
And Opus itself was already bad with biotech questions. The fact that they somehow made it WORSE for Fable is mindboggling.
Re: The classifiers Anthropic puts in front of Fable are too zealous
#78This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…
Re: The classifiers Anthropic puts in front of Fable are too zealous
#79Earlier quoted context omitted.
If your prompt has to do with those areas, yes. I haven't seen a single refusal yet. Reportedly the biology guiderails are particularly strict.
Fable refused to fix a Javascript error interfering with layout on our website. It's stupid and useless. It feels like whats really happening is Anthropic oversold Fable's claims; best case the CEO was given bad information; worst case they probably internally discovered it was cheating on benchmarks. Either case if feels like we're being lead on.
With these guardrails it is completely useless. The only hope is that they eventually convince the US Gov to let them use a saner classifier.
Re: The classifiers Anthropic puts in front of Fable are too zealous
#80Fable was refusing to patch vllm for me when trying to get mtp to work on r9700 gpus. Kept on bumping down to opus. Tried to really sanitize my prompts and everything but it seemed intrinsically prohibited from doing this sort of work. I guess it’s useful for making inane one shot games and websites, lol.