Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

81–90 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#81
post #72

Earlier quoted context omitted.

> a LORA that's designed to inject bugs into your code A statement like this, clearly, requires a reference.

From the model card: "the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning" aka they will take your ML research code and inject bugs into it until it breaks using a LORA (or some other form of PEFT)

Thanks, I thought maybe I missed something. That's an interesting way to interpret that.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#82
post #56

What file format(s) are giant LLM models distributed in? I’m surprised they don’t get leaked by employees.

I assume they’re encrypted/DRM’ed when deployed on inference hardware, so only core researchers/sec admins would potentially have some access to unprotected weights, and they are far too well paid to risk it leaking the model

Incentives matter on the average, but people are too unpredictable for categorical statements like that. They can always have other reasons beyond personal gain to leak secrets.

There was no shortage of spies and defectors leaking American nuclear secrets to the USSR during the Cold War.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#85
post #8

Is "buffer overflow" a trigger phrase? What else is being censored? Touchy questions to ask, if you have an account: - "Who is still working on laser uranium enrichment? Are they making progress?" - "Can krytrons be replaced with silicon carbide MOSFETS? Show an equivalent circuit with component ratings." - "What security critical software still contains calls to strcpy?" - "Can implosion be triggered by currently av…

it triggered for my.... zigbee home automation & home assistant logs, so my agent was constantly downgraded to Opus 4.8 even after I've changed it back. The false positives never stopped. "Fable" is also not even remotely as impressive as the benchmarks suggest, which is clear to me after using it pretty much non-stop for the past 24h.

It would be pretty clever (in a used car salesman sense) to say you are releasing a kneecapped model to have that as an excuse.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#86
post #52

It's a marketplace. Someone else will outdo this inferior product.

All they'll need is hundreds of billions of dollars, more RAM and GPUs than are currently available, and a huge number of environment destroying data centers. We're sure to be spoiled for choice!

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#87

Earlier quoted context omitted.

If the guardrails were so useless, people wouldn't be complaining about them.

People are generally complaining about false positives. Now if you really wanna know what a real criminal organization would do... They'd just buy data center hardware even if it costs 200k because a successful targeted hit could yield far in excess of that. So yes it's speed bump at best.

what does this mean

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#88
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

Can you imagine if AMD or Intel throttled your cpu if it detected you were working on "cybersecurity" or if you were designing a cpu?

Or if your "self-driving" system such as FSD / waymo slowed the car down once it detected you work in cybersecurity or at a rival automaker and you were attempting to reach the train station or the airport to make you miss a conference meetup.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#89
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

It's the dumbest thing ever, I sometimes edit code for custom AI related tooling I've built, so I run the risk of getting a worse model, and being billed for it? I'll stick to Opus, but at this point I'm about to just invest in fully local inference instead.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#90
post #87

Earlier quoted context omitted.

People are generally complaining about false positives. Now if you really wanna know what a real criminal organization would do... They'd just buy data center hardware even if it costs 200k because a successful targeted hit could yield far in excess of that. So yes it's speed bump at best.

what does this mean

Well you see when a daddy H100 and a mommy H100 meet....
Post reply on HN