Earlier quoted context omitted.
> a LORA that's designed to inject bugs into your code A statement like this, clearly, requires a reference.
From the model card: "the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning" aka they will take your ML research code and inject bugs into it until it breaks using a LORA (or some other form of PEFT)
Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
81–90 of 570 posts
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#82What file format(s) are giant LLM models distributed in? I’m surprised they don’t get leaked by employees.
I assume they’re encrypted/DRM’ed when deployed on inference hardware, so only core researchers/sec admins would potentially have some access to unprotected weights, and they are far too well paid to risk it leaking the model
There was no shortage of spies and defectors leaking American nuclear secrets to the USSR during the Cold War.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#83Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#84Somewhere I read that malware is already starting to use nuclear and biological and cybersecurity terms in the code to trick Fable into shutting down. Even if this is just a hypothetical attack vector so far, it seems likely to work.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#85Is "buffer overflow" a trigger phrase? What else is being censored? Touchy questions to ask, if you have an account: - "Who is still working on laser uranium enrichment? Are they making progress?" - "Can krytrons be replaced with silicon carbide MOSFETS? Show an equivalent circuit with component ratings." - "What security critical software still contains calls to strcpy?" - "Can implosion be triggered by currently av…
it triggered for my.... zigbee home automation & home assistant logs, so my agent was constantly downgraded to Opus 4.8 even after I've changed it back. The false positives never stopped. "Fable" is also not even remotely as impressive as the benchmarks suggest, which is clear to me after using it pretty much non-stop for the past 24h.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#86It's a marketplace. Someone else will outdo this inferior product.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#87Earlier quoted context omitted.
If the guardrails were so useless, people wouldn't be complaining about them.
People are generally complaining about false positives. Now if you really wanna know what a real criminal organization would do... They'd just buy data center hardware even if it costs 200k because a successful targeted hit could yield far in excess of that. So yes it's speed bump at best.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#88The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio
Can you imagine if AMD or Intel throttled your cpu if it detected you were working on "cybersecurity" or if you were designing a cpu?
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#89The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#90Earlier quoted context omitted.
People are generally complaining about false positives. Now if you really wanna know what a real criminal organization would do... They'd just buy data center hardware even if it costs 200k because a successful targeted hit could yield far in excess of that. So yes it's speed bump at best.
what does this mean