Live data from Hacker News

The classifiers Anthropic puts in front of Fable are too zealous

combine-lab.github.io

11–20 of 203 posts

Re: The classifiers Anthropic puts in front of Fable are too zealous

#12
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

I've had Fable read and write code. Never saw any downgrades.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#13
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

Anyone test the "Gay" jailbreak to see if it works on Fable?

Re: The classifiers Anthropic puts in front of Fable are too zealous

#14

This honestly just reads as “this model failed exactly where the company said it would but I’m very special and deserve special treatment rather than the same overactive guardrails I and everyone else were told we would get.”

When did Anthropic say you couldn't use it for math?

Re: The classifiers Anthropic puts in front of Fable are too zealous

#16
post #6

Do we think that someone at Anthropic, OpenAI, the government... has access to SOTA models without censorship? "How do I build an effective weapon?", "How do I effectively control the masses?"... It's very concerning that we get the nerfed models but you know that somewhere, people with a lot of resources have access to the raw, uncensored, probably more powerful models. The sprint toward AGI looks even more dangerou…

Yes. Mythos is almost exactly that. Willing to do in depth vulnerability and POC work.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#18
Fable was refusing to patch vllm for me when trying to get mtp to work on r9700 gpus. Kept on bumping down to opus. Tried to really sanitize my prompts and everything but it seemed intrinsically prohibited from doing this sort of work. I guess it’s useful for making inane one shot games and websites, lol.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#19
post #6

Do we think that someone at Anthropic, OpenAI, the government... has access to SOTA models without censorship? "How do I build an effective weapon?", "How do I effectively control the masses?"... It's very concerning that we get the nerfed models but you know that somewhere, people with a lot of resources have access to the raw, uncensored, probably more powerful models. The sprint toward AGI looks even more dangerou…

I think the problem here is that LLMs aren’t really “intelligence models” but more like “knowledge models”. LLMs don’t “think”, they just use a clever trick to make it seem like they do. I might not understand a lot about current state of AI, but that’s what they seem to be. Give it information and ask to organise it and make links, and they’ll do it, but that’s it, they don’t continually try to get out of the knowledge box they’re at, they don’t even know there’s a box.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#20
post #6

Do we think that someone at Anthropic, OpenAI, the government... has access to SOTA models without censorship? "How do I build an effective weapon?", "How do I effectively control the masses?"... It's very concerning that we get the nerfed models but you know that somewhere, people with a lot of resources have access to the raw, uncensored, probably more powerful models. The sprint toward AGI looks even more dangerou…

[deleted]
Post reply on HN