Earlier quoted context omitted.
Was this program available to independent security researchers or just established organizations? The docs you linked aren't very clear on this.
Any public research footprint seems to be enough, I applied as an individual and everyone I know who tried got accepted.
Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
71–80 of 570 posts
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#72Earlier quoted context omitted.
I've seen this claim a few times, but when I triggered the guardrails in Claude Code, it clearly notified me that it had switched to a different model ("something something for security purposes..."). Are you using Fable in Claude Code or in the browser?
They've said that they'll stop notifying developers when this gets triggered, instead they'll load in basically like a LORA that's designed to inject bugs into your code.
A statement like this, clearly, requires a reference.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#73Earlier quoted context omitted.
it triggered for my.... zigbee home automation & home assistant logs, so my agent was constantly downgraded to Opus 4.8 even after I've changed it back. The false positives never stopped. "Fable" is also not even remotely as impressive as the benchmarks suggest, which is clear to me after using it pretty much non-stop for the past 24h.
It has to be sort of impressive, given that you tried so hard to use it instead of the regular Opus.
I, for one, have tried using it several times today and the guardrails kept switching the model back to Opus, so I have no clue if it's impressive or not.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#74It's a marketplace. Someone else will outdo this inferior product.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#75Earlier quoted context omitted.
They've said that they'll stop notifying developers when this gets triggered, instead they'll load in basically like a LORA that's designed to inject bugs into your code.
> a LORA that's designed to inject bugs into your code A statement like this, clearly, requires a reference.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#76It's a marketplace. Someone else will outdo this inferior product.
And it doesn't look like OpenAI will have a good answer to Mythos anytime soon. Based on what their chief scientist wrote to staff recently (https://archive.is/fN2pg), GPT 5.6 is a "meaningful improvement" over 5.5 - in other words, just a normal version bump. And no news or even rumors regarding GPT 6.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#77I wonder how many millions they are wasting on putting up these guardrails when it's a completely useless exercise that is a speed bump at best.
If the guardrails were so useless, people wouldn't be complaining about them.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#78The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio
I don't.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#79The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio
Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price?
If the answer is no, could that be construed as fraud?