Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

51–60 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#51
For the last month, I've been making dramatic improvements to the security of the custom code developed at one of my customers using... GPT 5.5 dialed up to "Extra High" thinking.

It only pushes back sometimes if you ask it to create a "repro" that can be used to verify the vulnerability in production. Often it'll oblige, especially if you warn it not to create anything that could be actually harmful.

If the frontier models get locked down so that they flat refuse to do this kind of work, but Chinese and (less capable) open models aren't, then a lot of large enterprise orgs will be left twisting in the wind.

“AI can in principle help both the ‘good guys’ and the ‘bad guys’,” -- Dario Amodei

No Dario, no it can't, you've blocked one of those scenarios.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#53

This is a clickbait article with a garbage title. From the actual article, the one quoted cybersecurity researcher is sane about it: “But it is understandable as we are still in the early days and they are still adapting their guardrails. I am sure they are going to evolve over time as Anthropic and other frontier model companies will collaborate more with the current new generation of cybersecurity companies,” said…

I’m a cybersecurity researcher. Article seemed fine to me and echos a lot of me and my colleagues concerns. If you did regular malware analysis you would see that these groups already have access to LLMs that they’re using for development. What Anthropic is doing here is just hamstringing the good guys

I'm a cybersecurity researcher! Can you explain how Anthropic is just hamstringing the good guys?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#54

What file format(s) are giant LLM models distributed in? I’m surprised they don’t get leaked by employees.

What’s the point? Anthropic and other frontier vendors already provide their models on other services like vertex, bedrock, or openrouter It’s not like anyone can home lab one of these models without quite a bit of hardware

Yeah we can probably figure out how to run it on xiaomi gpus

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#55

Earlier quoted context omitted.

I’m a cybersecurity researcher. Article seemed fine to me and echos a lot of me and my colleagues concerns. If you did regular malware analysis you would see that these groups already have access to LLMs that they’re using for development. What Anthropic is doing here is just hamstringing the good guys

I'm a cybersecurity researcher! Can you explain how Anthropic is just hamstringing the good guys?

I did in my comment above.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#56

What file format(s) are giant LLM models distributed in? I’m surprised they don’t get leaked by employees.

I assume they’re encrypted/DRM’ed when deployed on inference hardware, so only core researchers/sec admins would potentially have some access to unprotected weights, and they are far too well paid to risk it leaking the model

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#57

Earlier quoted context omitted.

I'm a cybersecurity researcher! Can you explain how Anthropic is just hamstringing the good guys?

I did in my comment above.

You said these groups have access to LLMs. So what? Mythos/Fable are a step change above most LLMs. Responsibly limiting access and easing it up over time safely is the sane move.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#58
post #8

Is "buffer overflow" a trigger phrase? What else is being censored? Touchy questions to ask, if you have an account: - "Who is still working on laser uranium enrichment? Are they making progress?" - "Can krytrons be replaced with silicon carbide MOSFETS? Show an equivalent circuit with component ratings." - "What security critical software still contains calls to strcpy?" - "Can implosion be triggered by currently av…

it triggered for my.... zigbee home automation & home assistant logs, so my agent was constantly downgraded to Opus 4.8 even after I've changed it back. The false positives never stopped. "Fable" is also not even remotely as impressive as the benchmarks suggest, which is clear to me after using it pretty much non-stop for the past 24h.

It has to be sort of impressive, given that you tried so hard to use it instead of the regular Opus.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#60
post #15

Somewhere I read that malware is already starting to use nuclear and biological and cybersecurity terms in the code to trick Fable into shutting down. Even if this is just a hypothetical attack vector so far, it seems likely to work.

Confirmed: https://socket.dev/blog/mini-shai-hulud-miasma-and-hades-wor...
Post reply on HN