Earlier quoted context omitted.
The announcement elucidated this, and it's IMO worse than this. They don't downgrade to a cheaper model ([edit] for certain classes of offense they suspect you of). They sabotage the model's outputs in other, undisclosed, ways (specifically, "prompt modification, steering vectors, or parameter-efficient fine-tuning"). So, for example, they might load in a steering vector that just forgets the API to PyTorch. But it i…
It honestly explains so many issues I have been having, as I used it primarily for ML research (on my personal account, doing things not related to my job I should note). It would literally typo package names and spend huge amounts of time failing to setup simple environments…then do stupid things like set the learning rate to 1e-7, and use the eval set as training data.
Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
261–270 of 570 posts
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#262Earlier quoted context omitted.
Was this program available to independent security researchers or just established organizations? The docs you linked aren't very clear on this.
I was doing a CTF (with AI expected, even some anti-AI twists included) around the time the restrictions were tightened and was able to get approved by just saying it is a personal security research and doing a CTF. The experience was not nice though, it would happily chug away on a task and not even "hack this web", just asking about security of a binary was enough even with "this is a CTF handout..." - it would bur…
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#263Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#264Earlier quoted context omitted.
Yeah but a lot of the guardrails are pretty obviously to prevent competition not for safety.
Hmm. Maybe they are concerned about state actors trying to train equivalent models without the safeguards?
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#265These guardrails are solely a reason for using your data for training purposes. Every flagged message can be used for training.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#266Earlier quoted context omitted.
I’ve been working on a rather complex mapping project and have been getting MUCH better results with Fable than Opus.
So as not to be vague, and since I just pushed a version I'm starting to be vaguely happy with... https://tylereaves.github.io/uk-rail-map/ This is the result of probably a few hundred round trips. The really interesting part of the problem is keeping it both relatively true to real geometry, while greatly exaggerating it horizontally so you can actually see the individual running lines/sidings, like a signaling sche…
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#267Earlier quoted context omitted.
> it won't just reject ML research, which I can understand I don't.
They don't want someone to piggyback Anthropic's Mythos to make their own Mythos with less effort than it cost Anthropic.
I think if they want to behave anti competitively they should be honest about it and we should absolutely call them on it. Perhaps even regulators should.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#268If Claude Fable stops helping you, you'll never know
https://news.ycombinator.com/item?id=48467896
and Related:
Claude Fable 5
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#269Earlier quoted context omitted.
Can you imagine if AMD or Intel throttled your cpu if it detected you were working on "cybersecurity" or if you were designing a cpu?
Or if your "self-driving" system such as FSD / waymo slowed the car down once it detected you work in cybersecurity or at a rival automaker and you were attempting to reach the train station or the airport to make you miss a conference meetup.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#270The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio
> The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. My hypothesis is they know they can’t build effective enough guardrails, so scaring people into not trying is how they have decided to stop it.