I'm curious what the state of alignment research is. My gut says this is basically impossible. People have different moral frameworks. Each individual probably has an inconsistent moral framework. Even granting perfect consistency, applying these typically requires some knowledge of reality. And these LLM / harness combos are turing complete. So you don't know what it should do, you may not even know what you would d…
it's anthropics moral framework that matters, not the myriad of moral frameworks of the individual users
The classifiers Anthropic puts in front of Fable are too zealous
101–110 of 203 posts
Re: The classifiers Anthropic puts in front of Fable are too zealous
#102[flagged]
Re: The classifiers Anthropic puts in front of Fable are too zealous
#103I had one completely random trip when I was investigating some normal code. As far as I can tell a sub-agent ended up reading a file that tripped Fable during a review, but the whole feature was nowhere near anything secure so I don't know what could have caused it.
I also got completely locked out of Fable when working on parts of a subscription system (stripe subs).
But my experience isn't as bad as some peoples. The above maybe covers 15% of my attempted use cases. For the remaining 85% it has chugged along fine, sometimes in code I assumed would trigger it. It really feels random to me when it actually flags.
Re: The classifiers Anthropic puts in front of Fable are too zealous
#104Earlier quoted context omitted.
> It strikes me as highly unlikely that Anthropic has developed another fable-class model where the only difference is that it doesn't answer questions in that way I'm curious why you think that's highly unlikely given the monetary incentive (or even post-monetary!) to create such a thing? I imagine there's also an arms race aspect, if you assume your enemies (whoever they are) have access to such a model, certainly…
Cost, the politics of the people involved, and that there would be no real need for secrecy around it (but lots of need for marketing) so we'd probably know about it. It frankly doesn't seem like it would be that useful either... the US knows how to build weapons of mass destruction.
Re: The classifiers Anthropic puts in front of Fable are too zealous
#105This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…
I've been working with it heavily since its first release. I use it for software architects, complex debugging and some development and I have not had it refuse or downgrade even once.
Re: The classifiers Anthropic puts in front of Fable are too zealous
#106Re: The classifiers Anthropic puts in front of Fable are too zealous
#107Earlier quoted context omitted.
> the very thing Anthropic says it's not good for Where? Certainly not in its announcement, for one: https://platform.claude.com/docs/en/about-claude/models/intr... No "don't use this for X".
I thought it was the very first line of the product announcement, where they defined what it was they were calling "Fable" as opposed to "Mythos" in the first place: https://www.anthropic.com/news/claude-fable-5-mythos-5 > Today we’re launching Claude Fable 5: a Mythos-class model that we’ve made safe for general use. It then goes on to a lengthy and detailed section outlining the safety considerations: https://www.a…
Re: The classifiers Anthropic puts in front of Fable are too zealous
#108Earlier quoted context omitted.
The author is working on an opensource C++ codebase and not on biology tasks. The work is around tooling. It's like saying well a scalpel is used for medical reasons, sure. But manufacturing scalpels is metalworking, not medicine.
I think it's accurate to characterize the project as bio-related work: https://github.com/COMBINE-lab/salmon > salmon is a wicked-fast program for highly-accurate, transcript-level quantification from RNA-seq data. It pairs a fast mapping stage — selective alignment, or alignment-free sketch mode (--sketch) — with a massively-parallel statistical model (EM/VBEM over equivalence classes) to estimate transcript abundan…
Re: The classifiers Anthropic puts in front of Fable are too zealous
#109To summarize: the classifiers Anthropic puts in front of Fable are way, way too zealous and have way too many false positives. From my experience, the model itself is very useful when it isn't refusing any of your prompts.
And just to be clear, plan was already done, just had to review it, it got opus 4.8 Max and gpt 5.5 Extra High validated already and they didn't use much resource for it so I just don't get it. I guess they want to use it as a way to feed the extra credit money income.
I'm using a homemade ai consensus thing for planning and I wanted to add fable to it but forget it.
Or maybe I should use fable in low effort reasoning mode and it will be better than opus 4.8 at max ?
Re: The classifiers Anthropic puts in front of Fable are too zealous
#110Do we think that someone at Anthropic, OpenAI, the government... has access to SOTA models without censorship? "How do I build an effective weapon?", "How do I effectively control the masses?"... It's very concerning that we get the nerfed models but you know that somewhere, people with a lot of resources have access to the raw, uncensored, probably more powerful models. The sprint toward AGI looks even more dangerou…
"How do I build an effective weapon?" and "How do I effectively control the masses?" were research projects for the US government before you were born. One gave the world the Manhattan project, to name only one example, and the other MKULTRA. The government and cooperating companies had knowledge in both fields beyond the state of the art publicly available and continue to hold that edge over the public and foreign a…
One difference is that a CEO cannot set off an atomic bomb, but they can use an uncensored AGI model. The side-effects would be impossible to trace.
> I wonder, though, how many people advocating for popular access to uncensored AI models also advocate for an unrestricted (not infringed) right to bear arms or an unrestricted right to freedom of speech.
I advocate for all three of those things, for the same reason: the people I least want to have access to them, almost definitely do and it's imperative that the rest of us sit on equal footing.