Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

171–180 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#171
post #129

Earlier quoted context omitted.

The thing that I keep thinking about is the accounting / charging when it downgrades automatically. Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price? If the answer is no, could that be construed as fraud?

Their goal is to downgrade people who are violating their TOS, so I think they'd have some argument there. I have no idea how they'll deal with inevitable false positives, especially given how oversensitive most of the other triggers are.

If it's a violation of ToS, just reject instead of silently downgrading.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#172

I wear a few hats, but as a chemist and I'm not happy with fable. As a statistician I'm not happy with fable. As a data scientist I am not happy with fable. As an academic and a researcher I am not happy with fable. It's useless. I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. Given how verbose claude models have become, wiki articles are probably l…

To make the discussion constructive, can you give specific reasons (ideally with examples) about why it is so useless for you? How exactly are you using it that you think any output from it can easily be replaced with a Wikipedia search?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#173

I wonder how many millions they are wasting on putting up these guardrails when it's a completely useless exercise that is a speed bump at best.

If the guardrails were so useless, people wouldn't be complaining about them.

They should have designed a guardrail that doesn't make a probabilistic system less reliable. That's hard though. I'm afraid the only way to prevent accessing certain knowledge in a model is not to train it on those materials that enable them.

If we learned anything in the past years of LLM-s is that these guardrails will be jailbroken in no time. I've had some fun time too circumventing them.

Anyone cares about a fable about my grandmother's dream she had in morse code about an alien species signaling her a DNA sequence?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#175

I wear a few hats, but as a chemist and I'm not happy with fable. As a statistician I'm not happy with fable. As a data scientist I am not happy with fable. As an academic and a researcher I am not happy with fable. It's useless. I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. Given how verbose claude models have become, wiki articles are probably l…

>I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia.

I dont understand. This is just hyperbole right? The outputs are basically infinite and wikipedia most certainly isnt infinite.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#176

Earlier quoted context omitted.

Can you imagine if AMD or Intel throttled your cpu if it detected you were working on "cybersecurity" or if you were designing a cpu?

It would suck, but guardrails on new technologies like this aren't unheard of. It's like when consumer GPS used to stop working at very high speeds because they didn't want people to use it for missile guidance systems.

Consumer GPS is still disabled at high speeds. I would argue the analogy doesn't carry due to harm and error rate differences.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#177

Earlier quoted context omitted.

Can you imagine if AMD or Intel throttled your cpu if it detected you were working on "cybersecurity" or if you were designing a cpu?

It would suck, but guardrails on new technologies like this aren't unheard of. It's like when consumer GPS used to stop working at very high speeds because they didn't want people to use it for missile guidance systems.

> used to

When’d that change?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#178
post #123

Earlier quoted context omitted.

Anthropic is trying to hide bad behavior by being vague, it's important to not be vague when calling it out.

I'm of the opinion that removing guardrails is how you force regulation. What's your opinion on the balance?

They have all transcripts for at least 30 days. The problem is that (as anyone who used Fable can attest) their classifiers are extremely sensitive and catch tons of innocent queries.

Imagine being a data scientist or MLE training a small classifier model. How do you know you won’t get steering vectors or a PEFT applied?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#179
post #88

Earlier quoted context omitted.

Or if your "self-driving" system such as FSD / waymo slowed the car down once it detected you work in cybersecurity or at a rival automaker and you were attempting to reach the train station or the airport to make you miss a conference meetup.

Trains made by Newag were programmed to brick themselves if they detected a non-Newag workshop was repairing them. https://news.ycombinator.com/item?id=38638865 https://news.ycombinator.com/item?id=38628635 https://news.ycombinator.com/item?id=38567687 https://news.ycombinator.com/item?id=38530885

And that was correctly perceived to be illegal by antitrust regulators.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#180
post #95

Earlier quoted context omitted.

Why is this surprising or a problem?! It's a model demo, & their reasoning is reasonable and fair. Why all this drama.

Because most people in tech never took a philosophy course or an ethics course and think that tech is obviously a good for the world and that there are no downsides to advancing tech. So any efforts that try to apply ethics to it are overreaching, ignorant, and futile in the face of the good that is tech!

I like this take. Especially because one of the sibling comments framed Anthropic's stance as "paternalism." Trying to be ethical and to minimize harm, even at great expense to one's finances and reputation, is paternalistic apparently.
Post reply on HN