Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

211–220 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#211
post #15

Somewhere I read that malware is already starting to use nuclear and biological and cybersecurity terms in the code to trick Fable into shutting down. Even if this is just a hypothetical attack vector so far, it seems likely to work.

We all need to use nuclear, bio and cybersec terms in all our code to make low quality filtering like this untenable. When you can't work on a resume that has cybersecurity or biology terms in it or reply to a job opening that includes them because the "AI" filtering is so bad that it confuses these for threats, that deserves a collective response, particularly to an IPO'ing company that claims they'll make workers o…

That's why I use M-x spook to generate all of my variable names

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#212
post #162

I wear a few hats, but as a chemist and I'm not happy with fable. As a statistician I'm not happy with fable. As a data scientist I am not happy with fable. As an academic and a researcher I am not happy with fable. It's useless. I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. Given how verbose claude models have become, wiki articles are probably l…

I’ve been working on a rather complex mapping project and have been getting MUCH better results with Fable than Opus.

So as not to be vague, and since I just pushed a version I'm starting to be vaguely happy with...

https://tylereaves.github.io/uk-rail-map/

This is the result of probably a few hundred round trips. The really interesting part of the problem is keeping it both relatively true to real geometry, while greatly exaggerating it horizontally so you can actually see the individual running lines/sidings, like a signaling schematic.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#213

I wear a few hats, but as a chemist and I'm not happy with fable. As a statistician I'm not happy with fable. As a data scientist I am not happy with fable. As an academic and a researcher I am not happy with fable. It's useless. I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. Given how verbose claude models have become, wiki articles are probably l…

> Given how verbose claude models have become, wiki articles are probably less verbose too

Telling models to respond in the style of Wikipedia is one of the best ways to make their output bearable in my experience (for chat models, not agents)

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#214

Earlier quoted context omitted.

The thing that I keep thinking about is the accounting / charging when it downgrades automatically. Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price? If the answer is no, could that be construed as fraud?

The announcement elucidated this, and it's IMO worse than this. They don't downgrade to a cheaper model ([edit] for certain classes of offense they suspect you of). They sabotage the model's outputs in other, undisclosed, ways (specifically, "prompt modification, steering vectors, or parameter-efficient fine-tuning"). So, for example, they might load in a steering vector that just forgets the API to PyTorch. But it i…

It honestly explains so many issues I have been having, as I used it primarily for ML research (on my personal account, doing things not related to my job I should note). It would literally typo package names and spend huge amounts of time failing to setup simple environments…then do stupid things like set the learning rate to 1e-7, and use the eval set as training data.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#215

I'd like to offer a counter-point to many of the comments here. While I understand being stymied and frustrated by a product one is paying for... At the same time, I personally think the tradeoff between "having guardrails" and "some users are unhappy with the product" is well worth it. Think of what would happen if all of us who aren't so well intentioned could exploit Fable in terrible ways. Surely this tradeoff is…

Yeah but a lot of the guardrails are pretty obviously to prevent competition not for safety.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#216

I wear a few hats, but as a chemist and I'm not happy with fable. As a statistician I'm not happy with fable. As a data scientist I am not happy with fable. As an academic and a researcher I am not happy with fable. It's useless. I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. Given how verbose claude models have become, wiki articles are probably l…

To make the discussion constructive, can you give specific reasons (ideally with examples) about why it is so useless for you? How exactly are you using it that you think any output from it can easily be replaced with a Wikipedia search?

The cybersecurity and bioweapons filters reach so far that they set in as soon as the model even glazes anything STEM-related. It might give a good impression of ones ex or write a decent fanfiction but anything that could bring humanity forward is strictly off-limits.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#217

I'd like to offer a counter-point to many of the comments here. While I understand being stymied and frustrated by a product one is paying for... At the same time, I personally think the tradeoff between "having guardrails" and "some users are unhappy with the product" is well worth it. Think of what would happen if all of us who aren't so well intentioned could exploit Fable in terrible ways. Surely this tradeoff is…

Yeah but a lot of the guardrails are pretty obviously to prevent competition not for safety.

Hmm. Maybe they are concerned about state actors trying to train equivalent models without the safeguards?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#218
The guardrails are pretty tight. It is even refusing to decode morse code: https://x.com/Schappi/status/2064839631137546503?s=20

The prompt was: please translate .. ..-. / -.-- --- ..- / -.-. .- -. / .-. . .- -.. / - .... .. ... --..-- / - --- ..- -.-. .... / --. .-. .- ... ...

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#219

I'd like to offer a counter-point to many of the comments here. While I understand being stymied and frustrated by a product one is paying for... At the same time, I personally think the tradeoff between "having guardrails" and "some users are unhappy with the product" is well worth it. Think of what would happen if all of us who aren't so well intentioned could exploit Fable in terrible ways. Surely this tradeoff is…

What would happen, exactly?

My imagination says “nothing much”.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#220
post #169
post #121

Earlier quoted context omitted.

> We will require 30-day retention for all traffic on Mythos-class models, on both first- and third-party surfaces. We won’t use this data to train new Claude models, or for any non-safety-related purpose Whatever problem we might have with them, they explicitly say that they do not do this in the launch post.

"We won’t use this data to train new Claude models" What about non-Claude models?

"Introducing our latest model, CIaude, spelled with a capital "i" and legally distinct from Claude!"
Post reply on HN