Live data from Hacker News

The classifiers Anthropic puts in front of Fable are too zealous

combine-lab.github.io

151–160 of 203 posts

Re: The classifiers Anthropic puts in front of Fable are too zealous

#151

Earlier quoted context omitted.

I used to agree with you but overtime I’ve changed my mind. For reference I created predictive linguistics at Google in the first products and this is a many order scale up of that, with new complexities of course. The best analogy I can give you is that it is a really advanced synthesis machine, which looks like human thought but is more of a hyper advance “replay” of human thought in various contexts. Where you beg…

I feel like that's more saying they can't train on the fly, and also that serializing spatial data and world models is something we haven't really done fully. For me all neural networks synthetic or otherwise are replay machines or stream prediction machines. Nerve signals in, and nerve signals out. If I create output signals to the muscle nerves like this when my eyes see signals like that, good things happen, i get…

I’ll Disagree and all I’ll say is when you say messier that hand waving away the differences between an RC car and a real car because they both drive but the real car just has some messier complications.

That messier part is the complexity that is the difference.

What we have is a model. It’s still very distant from the original in meaningful ways.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#152

Bottom line: “California AI” (in Yann LeCun’s terminology) can not be relied upon. It could change at any time, and stop working for your project. For the future of AI, we need to look elsewhere.

Would be nice to have a distributed, independent AIs, each being trained their own way. Maybe it would have to be a really slow training process to keep costs low (years even?).

Re: The classifiers Anthropic puts in front of Fable are too zealous

#153
post #64
post #46

Earlier quoted context omitted.

It feels like the longtermist believers got involved in this (those are the people obsessed with garage-engineered designer viruses who have a very tenuous grasp on how biology research actually works).

Yeah i'm wondering how much of a role that plays in this as well. On the one hand I could believe it's something more benign, or the usual misunderstood fear mongering making it to some political level (well make sure those users can't get online anonymously! being our current craze). That said, chemistry and to some level physics have been the major domain of limited knowledge (chemistry because the average person c…

Nobody has tried to limit knowledge of chemistry or physics unless it was directly about doing something illegal, to the point of basically being a detailed recipe. Usually not even then. And when they have tried they've had basically zero success.

The ability for a handful of companies, simultaneously very powerful and easily susceptible to pressure from other powerful actors, to do the same sort of thing with the next generation of core learning and engineering tools, is freaking terrifying.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#154
post #126

Earlier quoted context omitted.

Were your question in math areas related to ML? They also restrict model development and research pretty heavily.

No. Robust control theory was one case. Dynamical systems another.

Well... now you know one direction Anthropic is looking towards for the future research.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#155
post #15

The honest way to say this is that Fable is not useful for bio-related work. The author is working on processing RNA sequences and similar biology tasks, and Fable's classifier has a hair trigger on those tasks.

> The honest way to say this is that Fable is not useful for bio-related work.

It is way worse than that. Try "How does digestion work?" and you will see "Fable's safeguards flagged this message". It's a stupid rate of false positives.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#156
post #5

This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…

And biology is by far the classifier's least favorite topic. It's not even close. I've had it downgrade to Opus for the following questions: "How confident are we that English and American Eels both spawn in the Sargasso Sea?" "Come up with five Zoology questions of increasing difficulty for a trivia game." "What's your favorite sarcopterygian?" My wife has some zoology-related preferences in her user instructions, a…

It would not even help me with updating my CV because I work in biology...

Re: The classifiers Anthropic puts in front of Fable are too zealous

#157
post #128

I am wondering about the author's allegation that there is a user filter, not just a prompt filter. Of course it could also be the case that it is just a prompt filter, but Fable sees memories from the authors' prior sessions that cause a rejection. I wonder if the author could control for this is in some way, if Claude lets you run isolated session without memory access.

The first rejection and subsequent modifications trying to adjust the prompt to pass the classifier might have gotten the account flagged so the classifier is now set to 'hair trigger'. While I'm not aware of Anthropic admitting they put flagged accounts in classifier 'jail', they previously showed they're aware how vulnerable any LLM is to jailbreaking with the 'silent switch' to 4.8, whose only purpose was to remove feedback signals from iterative jailbreak testing.

The obvious failure mode is that trying to fix an innocent prompt to pass an over-sensitive classifier looks like a bad actor trying to jailbreak the model. I don't really see how Anthropic can fix this. Jailbreaking is a fundamental weakness endemic to LLMs, so 'smarter' models aren't the answer.

I suspect they're being so stringent because, at least some at Anthropic, genuinely believe LLMs are already an existential risk to humanity. However, it's clear other frontier competitors rank that risk lower and are taking a more nuanced, pragmatic position on safety. To the extent Anthropic's fears continue to make them less useful to customers, competitors are going to bypass them.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#158
post #46

Earlier quoted context omitted.

And biology is by far the classifier's least favorite topic. It's not even close. I've had it downgrade to Opus for the following questions: "How confident are we that English and American Eels both spawn in the Sargasso Sea?" "Come up with five Zoology questions of increasing difficulty for a trivia game." "What's your favorite sarcopterygian?" My wife has some zoology-related preferences in her user instructions, a…

It feels like the longtermist believers got involved in this (those are the people obsessed with garage-engineered designer viruses who have a very tenuous grasp on how biology research actually works).

No, by far the most parsimonious explanation is they got slapped by a capricious US government so they went overboard on caution in an attempt not to generate any more controversy. A predictable response of chaotic government regulation.

Re: The classifiers Anthropic puts in front of Fable are too zealous

#159
post #134
post #128

I am wondering about the author's allegation that there is a user filter, not just a prompt filter. Of course it could also be the case that it is just a prompt filter, but Fable sees memories from the authors' prior sessions that cause a rejection. I wonder if the author could control for this is in some way, if Claude lets you run isolated session without memory access.

Author mentions trying incognito mode without success. I also tried their strictly mathematical problem description and got filtered 5/5 times.

That's interesting. I assumed that the OP's attempts to fix the prompt looked like jailbreaking attempts and got the account auto-flagged into hair-trigger 'classifier jail'. Of course, a bad actor would swap accounts, so maybe Anthropic flags both the account and the prompt (coming from any account).

Re: The classifiers Anthropic puts in front of Fable are too zealous

#160
Yeah, ran into this. I asked it to review a server I wrote for security vulnerabilities and it was "flagged" (after spending some money, of course). Kind of bizarre, there were so many ways a person could look at this and tell it was legit: the git log (look at my git config vs the author email and notice they're the same), the fact that none of this code is on the internet (private repo), the phrasing of my request, the fact that there's a long history of me collaborating with claude on building this, etc. I know someone's going to say: "the governments fault!" Yeah, to a point, but this wouldn't be an issue if these guys weren't relentlessly doom trolling or pretending like we're in a race with china. (What race exactly? To see who can enshittify the internet the fastest?) I wouldn't say I'm particularly upset about this, because before I tried it I had already read how other models have been able to find the same class of bugs, so I was using it more out of curiosity than need, but it does reinforce that these companies can take away these tools on a whim. Also, I just can't help but think that if your PR and marketing is literally making your software illegal to use, and causing people to hate you, maybe you're not doing it right.
Post reply on HN