Live data from Hacker News

The classifiers Anthropic puts in front of Fable are too zealous

combine-lab.github.io

101–110 of 203 posts

Re: The classifiers Anthropic puts in front of Fable are too zealous

#101
post #50

I'm curious what the state of alignment research is. My gut says this is basically impossible. People have different moral frameworks. Each individual probably has an inconsistent moral framework. Even granting perfect consistency, applying these typically requires some knowledge of reality. And these LLM / harness combos are turing complete. So you don't know what it should do, you may not even know what you would d…

it's anthropics moral framework that matters, not the myriad of moral frameworks of the individual users

Yeah and Anthropic is a... dividual consisting of founders, staff, and shareholders, and must comply with various governments ultimately deriving their values from billions of people.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#103
I've had mixed results with downgrading on Fable. I was able to do a complete audit of my OAuth implementation without any issue. But when I asked for an OWASP top-ten review of my code base it got through 5 of 6 tasks and tripped in the final summary, which Opus had to finish.

I had one completely random trip when I was investigating some normal code. As far as I can tell a sub-agent ended up reading a file that tripped Fable during a review, but the whole feature was nowhere near anything secure so I don't know what could have caused it.

I also got completely locked out of Fable when working on parts of a subscription system (stripe subs).

But my experience isn't as bad as some peoples. The above maybe covers 15% of my attempted use cases. For the remaining 85% it has chugged along fine, sometimes in code I assumed would trigger it. It really feels random to me when it actually flags.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#104
post #71
post #65

Earlier quoted context omitted.

> It strikes me as highly unlikely that Anthropic has developed another fable-class model where the only difference is that it doesn't answer questions in that way I'm curious why you think that's highly unlikely given the monetary incentive (or even post-monetary!) to create such a thing? I imagine there's also an arms race aspect, if you assume your enemies (whoever they are) have access to such a model, certainly…

Cost, the politics of the people involved, and that there would be no real need for secrecy around it (but lots of need for marketing) so we'd probably know about it. It frankly doesn't seem like it would be that useful either... the US knows how to build weapons of mass destruction.

I have to disagree. For the sake of argument say Elon Musk had his own personal, uncensored SOTA model. He has the cash and politics to make that a realistic goal. Would people want him to have that? Not really, hence secrecy as well.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#105
post #81
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

I've been working with it heavily since its first release. I use it for software architects, complex debugging and some development and I have not had it refuse or downgrade even once.

Daily use here, about 2.5 weekly 20x limits, never got flagged for code topics including finding memory safety vulnerabilities in my C++ project, but just got flagged for the first time for biology-related topics because I asked it to implement crop genetics and cross breeding into my game. Was able to bypass it by having opus reword the prompt (gene -> trait, cross breed -> trait mixing), and, critically, insisted that it not use any biology related words in its thinking or responses.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#107
post #89

Earlier quoted context omitted.

> the very thing Anthropic says it's not good for Where? Certainly not in its announcement, for one: https://platform.claude.com/docs/en/about-claude/models/intr... No "don't use this for X".

I thought it was the very first line of the product announcement, where they defined what it was they were calling "Fable" as opposed to "Mythos" in the first place: https://www.anthropic.com/news/claude-fable-5-mythos-5 > Today we’re launching Claude Fable 5: a Mythos-class model that we’ve made safe for general use. It then goes on to a lengthy and detailed section outlining the safety considerations: https://www.a…

It's not detailed, at all. It's all unnecessary verbiage and some meaningless graphs around "trust us, only we know what is safe". Meanwhile people run into these stupid "safeguards" on the most innocent queries. See e.g. this thread of discussion: https://news.ycombinator.com/item?id=48837404 Or indeed the very article these comments are under

Re: The classifiers Anthropic puts in front of Fable are too zealous

#108
post #24

Earlier quoted context omitted.

The author is working on an opensource C++ codebase and not on biology tasks. The work is around tooling. It's like saying well a scalpel is used for medical reasons, sure. But manufacturing scalpels is metalworking, not medicine.

I think it's accurate to characterize the project as bio-related work: https://github.com/COMBINE-lab/salmon > salmon is a wicked-fast program for highly-accurate, transcript-level quantification from RNA-seq data. It pairs a fast mapping stage — selective alignment, or alignment-free sketch mode (--sketch) — with a massively-parallel statistical model (EM/VBEM over equivalence classes) to estimate transcript abundan…

It's also data that Anthropic likely scraped and included in their training data.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#109

To summarize: the classifiers Anthropic puts in front of Fable are way, way too zealous and have way too many false positives. From my experience, the model itself is very useful when it isn't refusing any of your prompts.

It is, but it's also using tokens at absurd rate, I asked it to review the planned architecture for a medium scale project and it used my 5 hours limit on one prompt just zaaaaap, not even the fable limit straight up the full 5 hour session no more Claude for the afternoon thank you for paying you Max x20 sub. Hell it didn't even bother to finish produce anything worthwhile.

And just to be clear, plan was already done, just had to review it, it got opus 4.8 Max and gpt 5.5 Extra High validated already and they didn't use much resource for it so I just don't get it. I guess they want to use it as a way to feed the extra credit money income.

I'm using a homemade ai consensus thing for planning and I wanted to add fable to it but forget it.

Or maybe I should use fable in low effort reasoning mode and it will be better than opus 4.8 at max ?

Re: The classifiers Anthropic puts in front of Fable are too zealous

#110
post #6

Do we think that someone at Anthropic, OpenAI, the government... has access to SOTA models without censorship? "How do I build an effective weapon?", "How do I effectively control the masses?"... It's very concerning that we get the nerfed models but you know that somewhere, people with a lot of resources have access to the raw, uncensored, probably more powerful models. The sprint toward AGI looks even more dangerou…

"How do I build an effective weapon?" and "How do I effectively control the masses?" were research projects for the US government before you were born. One gave the world the Manhattan project, to name only one example, and the other MKULTRA. The government and cooperating companies had knowledge in both fields beyond the state of the art publicly available and continue to hold that edge over the public and foreign a…

The prior art you state is exactly why I think this is almost certainly happening.

One difference is that a CEO cannot set off an atomic bomb, but they can use an uncensored AGI model. The side-effects would be impossible to trace.

> I wonder, though, how many people advocating for popular access to uncensored AI models also advocate for an unrestricted (not infringed) right to bear arms or an unrestricted right to freedom of speech.

I advocate for all three of those things, for the same reason: the people I least want to have access to them, almost definitely do and it's imperative that the rest of us sit on equal footing.

Post reply on HN