Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

261–270 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#261

Earlier quoted context omitted.

The announcement elucidated this, and it's IMO worse than this. They don't downgrade to a cheaper model ([edit] for certain classes of offense they suspect you of). They sabotage the model's outputs in other, undisclosed, ways (specifically, "prompt modification, steering vectors, or parameter-efficient fine-tuning"). So, for example, they might load in a steering vector that just forgets the API to PyTorch. But it i…

It honestly explains so many issues I have been having, as I used it primarily for ML research (on my personal account, doing things not related to my job I should note). It would literally typo package names and spend huge amounts of time failing to setup simple environments…then do stupid things like set the learning rate to 1e-7, and use the eval set as training data.

That’s insane. I hope they fix it.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#262

Earlier quoted context omitted.

Was this program available to independent security researchers or just established organizations? The docs you linked aren't very clear on this.

I was doing a CTF (with AI expected, even some anti-AI twists included) around the time the restrictions were tightened and was able to get approved by just saying it is a personal security research and doing a CTF. The experience was not nice though, it would happily chug away on a task and not even "hack this web", just asking about security of a binary was enough even with "this is a CTF handout..." - it would bur…

[dead]

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#263

Earlier quoted context omitted.

You've already explicitly enabled extra usage in your account settings though, it is not on by default

Unknowingly. Is that set at the org level? Because I never set it and never had it do that before.

It is at the org level

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#264

Earlier quoted context omitted.

Yeah but a lot of the guardrails are pretty obviously to prevent competition not for safety.

Hmm. Maybe they are concerned about state actors trying to train equivalent models without the safeguards?

If a for profit company does a thing that could be motivated by profit or altruism, which of those 2 motivations do you think is most likely?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#265

These guardrails are solely a reason for using your data for training purposes. Every flagged message can be used for training.

If we're doing conspiracy theories what if fable is really dumb and not better than opus and the guardrails hide that nicely. Meanwhile the hype train keeps chugging.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#266
post #212
post #162

Earlier quoted context omitted.

I’ve been working on a rather complex mapping project and have been getting MUCH better results with Fable than Opus.

So as not to be vague, and since I just pushed a version I'm starting to be vaguely happy with... https://tylereaves.github.io/uk-rail-map/ This is the result of probably a few hundred round trips. The really interesting part of the problem is keeping it both relatively true to real geometry, while greatly exaggerating it horizontally so you can actually see the individual running lines/sidings, like a signaling sche…

Fascinating. Can you explain why southern London is DC while northern London is AC?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#267

Earlier quoted context omitted.

> it won't just reject ML research, which I can understand I don't.

They don't want someone to piggyback Anthropic's Mythos to make their own Mythos with less effort than it cost Anthropic.

So they are lying then when they say it's for safety reasons.

I think if they want to behave anti competitively they should be honest about it and we should absolutely call them on it. Perhaps even regulators should.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#269
post #88

Earlier quoted context omitted.

Can you imagine if AMD or Intel throttled your cpu if it detected you were working on "cybersecurity" or if you were designing a cpu?

Or if your "self-driving" system such as FSD / waymo slowed the car down once it detected you work in cybersecurity or at a rival automaker and you were attempting to reach the train station or the airport to make you miss a conference meetup.

Didn’t uber catch a lot of shit for nerfing the app for people suspected to be enforcing the laws they were breaking?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#270
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

> The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. My hypothesis is they know they can’t build effective enough guardrails, so scaring people into not trying is how they have decided to stop it.

[deleted]
Post reply on HN