Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

441–450 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#441
post #414

Earlier quoted context omitted.

Corporate America never backs down. It simply rallies and tries again later until people are too fatigued to care. The only solution is to abandon ship, which I am doing. MS walked back in OS ads the first few times, but ultimately we still ended up on the exact trajectory everyone was outraged at. OpenAI still ended up on its path to closed AI despite initial walk backs. The story repeats itself over and over again,…

Same with VISA/Mastercard deciding what we can/cannot buy. The only solution is to stop using their credit cards at all.

Yes, Monero is a lot better than credit cards for privacy and freedom. I hope to see it accepted more.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#442
post #129

Earlier quoted context omitted.

The thing that I keep thinking about is the accounting / charging when it downgrades automatically. Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price? If the answer is no, could that be construed as fraud?

Their goal is to downgrade people who are violating their TOS, so I think they'd have some argument there. I have no idea how they'll deal with inevitable false positives, especially given how oversensitive most of the other triggers are.

Their detection is too aggressive. Just today I'm trying to build a kernel for some SBC and I hit that downgrade. I just asked some things about `make menuconfig` items. I suppose it just flags everything related to linux kernel as cyber attacks.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#444

Wait a few months and a competitor will release a similarly powerful model with less guardrails, if they steal sufficient market share Anthropic will reverse policies. This is why I’m immensely hoping the Chinese don’t stop with their open sourced local models. None of these companies are your friend.

> This is why I’m immensely hoping the Chinese don’t stop with their open sourced local models. None of these companies are your friend.

The Chinese aren't your friend either [1].

[1] https://www.hks.harvard.edu/centers/carr-ryan/our-work/carr-...

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#445
post #417
post #395

Earlier quoted context omitted.

Yes, that's true. Excluding Fable, OAI models are the most refusal heavy. However, I'd rather get a refusal than response with poisoned output. Since currently there's no way to verify if poisoning happened or not, I don't trust Anthropic anymore, regardless of what they say. But my trust towards OAI is also brittle - what if they also do it, or start doing it? I want to have a verifiable way to know that the prompt…

What kind of work are you getting refusals on? Genuinely curious. The only refusal I’ve had in recent memory was declining to find doorbell camera footage matching a certain description, which is fair enough and I think EU laws heavily restrict such activities (even tho I’m not in the EU)

How would the AI be able to find the footage itself?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#446
post #241

Earlier quoted context omitted.

How does it help?

By withholding it from bad actors.

It withholds it from good actors (they cannot use it to harden their code against bad actors) and assumes bad actors don't have access to such tools anyway.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#448

Earlier quoted context omitted.

Wow… just wow. The future looks incredibly bleak if people are throwing fisftuls of money at this company. Anthropic will quickly become the sole arbiter of everything in your life.

Why do people think this is the future? Anthropic has the leading model, and so they're able to hold back functionality. They do so with obvious regards to safety. If anything a future with models of such capabilities and no safeguards would be a bleak future. But its likely what were headed in once other companies catch up.

I think it’s safe to say that many of us feel a lot less safe directly because of these policies and the inferred intentions of the company behind them. Nobody is arguing for unsafe models. We just don’t want to live in the plot of Deus Ex.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#449
Malware authors are pretty excited about guard-rails. you can add prompts to your malware to get LLM scanners to hit guard-rails and stop their runs. New shai-hulud npm worm campaign for example includes prompts to request biological weapon schematics/creation etc. to ensure LLM scanners probing NPM packages refuse to scan it.

These AI places have 0 clue about how threat actors actually work. None of their mitigations or guard-rails is effective, and now they are even turned against them.

Additionally, if they don't all implement the same level of effective guard-rails, there will always be some model you can abuse to do the work anyway, and hence there is 0 effect on threat actors, they will just run some local model that does 5% less quality, which does not matter to them 1 bit.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#450

Earlier quoted context omitted.

They don't want someone to piggyback Anthropic's Mythos to make their own Mythos with less effort than it cost Anthropic.

Ironic, given they piggybacked on the entirety of human knowledge and massive amounts of GPL'd software and repeatedly say they want to replace people with a tool. And now they say that's fine so long as people are entertained.

Pulling up the ladder behind you is a tradition as old as time.
Post reply on HN