Somewhere I read that malware is already starting to use nuclear and biological and cybersecurity terms in the code to trick Fable into shutting down. Even if this is just a hypothetical attack vector so far, it seems likely to work.
We all need to use nuclear, bio and cybersec terms in all our code to make low quality filtering like this untenable. When you can't work on a resume that has cybersecurity or biology terms in it or reply to a job opening that includes them because the "AI" filtering is so bad that it confuses these for threats, that deserves a collective response, particularly to an IPO'ing company that claims they'll make workers o…
Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
211–220 of 570 posts
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#212I wear a few hats, but as a chemist and I'm not happy with fable. As a statistician I'm not happy with fable. As a data scientist I am not happy with fable. As an academic and a researcher I am not happy with fable. It's useless. I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. Given how verbose claude models have become, wiki articles are probably l…
I’ve been working on a rather complex mapping project and have been getting MUCH better results with Fable than Opus.
https://tylereaves.github.io/uk-rail-map/
This is the result of probably a few hundred round trips. The really interesting part of the problem is keeping it both relatively true to real geometry, while greatly exaggerating it horizontally so you can actually see the individual running lines/sidings, like a signaling schematic.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#213I wear a few hats, but as a chemist and I'm not happy with fable. As a statistician I'm not happy with fable. As a data scientist I am not happy with fable. As an academic and a researcher I am not happy with fable. It's useless. I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. Given how verbose claude models have become, wiki articles are probably l…
Telling models to respond in the style of Wikipedia is one of the best ways to make their output bearable in my experience (for chat models, not agents)
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#214Earlier quoted context omitted.
The thing that I keep thinking about is the accounting / charging when it downgrades automatically. Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price? If the answer is no, could that be construed as fraud?
The announcement elucidated this, and it's IMO worse than this. They don't downgrade to a cheaper model ([edit] for certain classes of offense they suspect you of). They sabotage the model's outputs in other, undisclosed, ways (specifically, "prompt modification, steering vectors, or parameter-efficient fine-tuning"). So, for example, they might load in a steering vector that just forgets the API to PyTorch. But it i…
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#215I'd like to offer a counter-point to many of the comments here. While I understand being stymied and frustrated by a product one is paying for... At the same time, I personally think the tradeoff between "having guardrails" and "some users are unhappy with the product" is well worth it. Think of what would happen if all of us who aren't so well intentioned could exploit Fable in terrible ways. Surely this tradeoff is…
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#216I wear a few hats, but as a chemist and I'm not happy with fable. As a statistician I'm not happy with fable. As a data scientist I am not happy with fable. As an academic and a researcher I am not happy with fable. It's useless. I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. Given how verbose claude models have become, wiki articles are probably l…
To make the discussion constructive, can you give specific reasons (ideally with examples) about why it is so useless for you? How exactly are you using it that you think any output from it can easily be replaced with a Wikipedia search?
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#217I'd like to offer a counter-point to many of the comments here. While I understand being stymied and frustrated by a product one is paying for... At the same time, I personally think the tradeoff between "having guardrails" and "some users are unhappy with the product" is well worth it. Think of what would happen if all of us who aren't so well intentioned could exploit Fable in terrible ways. Surely this tradeoff is…
Yeah but a lot of the guardrails are pretty obviously to prevent competition not for safety.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#218The prompt was: please translate .. ..-. / -.-- --- ..- / -.-. .- -. / .-. . .- -.. / - .... .. ... --..-- / - --- ..- -.-. .... / --. .-. .- ... ...
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#219I'd like to offer a counter-point to many of the comments here. While I understand being stymied and frustrated by a product one is paying for... At the same time, I personally think the tradeoff between "having guardrails" and "some users are unhappy with the product" is well worth it. Think of what would happen if all of us who aren't so well intentioned could exploit Fable in terrible ways. Surely this tradeoff is…
My imagination says “nothing much”.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#220Earlier quoted context omitted.
> We will require 30-day retention for all traffic on Mythos-class models, on both first- and third-party surfaces. We won’t use this data to train new Claude models, or for any non-safety-related purpose Whatever problem we might have with them, they explicitly say that they do not do this in the launch post.
"We won’t use this data to train new Claude models" What about non-Claude models?