Live data from Hacker News

The classifiers Anthropic puts in front of Fable are too zealous

combine-lab.github.io

21–30 of 203 posts

Re: The classifiers Anthropic puts in front of Fable are too zealous

#22
post #3

Terrible title. Should be "Fable's guard rails are way too sensitive", which I don't think you can really blame Anthropic for. They likely had to whack them way up so it would block whatever trivial stuff got demoed to the government. I would expect them to dial down the sensitivity in a few months when nobody is looking.

>which I don't think you can really blame Anthropic for.

on the contrary, you can, and you should. their greasy effective altruist had always been by far the loudest proponent of the `safety` theater.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#24
post #15

The honest way to say this is that Fable is not useful for bio-related work. The author is working on processing RNA sequences and similar biology tasks, and Fable's classifier has a hair trigger on those tasks.

The author is working on an opensource C++ codebase and not on biology tasks. The work is around tooling.

It's like saying well a scalpel is used for medical reasons, sure. But manufacturing scalpels is metalworking, not medicine.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#25
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

If your prompt has to do with those areas, yes. I haven't seen a single refusal yet. Reportedly the biology guiderails are particularly strict.

I've had it refuse to help build an image classifier ml pipeline, pretty innocuous stuff. Got around it eventually but still it's a very dumb constraint to add to an otherwise very smart system

Re: The classifiers Anthropic puts in front of Fable are too zealous

#26
post #8

I have only really used Fable as a final pass on something. A "Take a look at everything we did so far, and make sure we didn't forget something" kind of review prompt. But it is a huge waste of money for most coding tasks. Opus is still overkill most of the time, too.

> But it is a huge waste of money for most coding tasks.

The key is not to indiscriminately use the most powerful/expensive model you can for everything. When you use it for what it's uniquely suited for and ask it to spawn subagents using Opus and Sonnet based on what tasks need, you'll get better results at a reasonable cost.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#27
post #6

Do we think that someone at Anthropic, OpenAI, the government... has access to SOTA models without censorship? "How do I build an effective weapon?", "How do I effectively control the masses?"... It's very concerning that we get the nerfed models but you know that somewhere, people with a lot of resources have access to the raw, uncensored, probably more powerful models. The sprint toward AGI looks even more dangerou…

Yes, of course.

Until Fable even the public had practically uncensored access to SOTA anthropic models (there were classifiers - but they were very hard to hit). And I'd have to double check but I'm pretty certain the public still has uncensored access to SOTA models from google (via GCP under threat of Google ceasing to do business with you and theoretically suing you if you violate the TOS).

Censorship being what they are doing here - preventing you from accessing the model for certain tasks. Censorship not being what a bunch of... motivated people... have been incorrectly suggesting is censorship: developing models to give the kinds of answers that the model developers want them to give - which has generally been a model that gives responses appropriate for a non-pornographic non-military business environment.

It strikes me as highly unlikely that Anthropic has developed another fable-class model where the only difference is that it doesn't answer questions in that way - e.g. that they have a fable model fine tuned so that when you ask it to develop biological weapons it responds similarly to asking fable to develop 3d rendering software. Of course, with uncensored access to the model it is likely possible to prompt it to develop biological weapons despite its inclination to decline.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#28
post #13
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

Anyone test the "Gay" jailbreak to see if it works on Fable?

That wasn't even effective on ChatGPT because the results were not detailed enough, at least with Meth, in my very short testing and based on the examples.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#29
post #6

Do we think that someone at Anthropic, OpenAI, the government... has access to SOTA models without censorship? "How do I build an effective weapon?", "How do I effectively control the masses?"... It's very concerning that we get the nerfed models but you know that somewhere, people with a lot of resources have access to the raw, uncensored, probably more powerful models. The sprint toward AGI looks even more dangerou…

I think the problem here is that LLMs aren’t really “intelligence models” but more like “knowledge models”. LLMs don’t “think”, they just use a clever trick to make it seem like they do. I might not understand a lot about current state of AI, but that’s what they seem to be. Give it information and ask to organise it and make links, and they’ll do it, but that’s it, they don’t continually try to get out of the knowle…

Feels like a distinction without a difference. What is any intelligence but a sum of its knowledge?

Re: The classifiers Anthropic puts in front of Fable are too zealous

#30
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

If your prompt has to do with those areas, yes. I haven't seen a single refusal yet. Reportedly the biology guiderails are particularly strict.

Fable refused to fix a Javascript error interfering with layout on our website.

It's stupid and useless.

It feels like whats really happening is Anthropic oversold Fable's claims; best case the CEO was given bad information; worst case they probably internally discovered it was cheating on benchmarks. Either case if feels like we're being lead on.

Post reply on HN