Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

471–480 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#471
post #417
post #395

Earlier quoted context omitted.

Yes, that's true. Excluding Fable, OAI models are the most refusal heavy. However, I'd rather get a refusal than response with poisoned output. Since currently there's no way to verify if poisoning happened or not, I don't trust Anthropic anymore, regardless of what they say. But my trust towards OAI is also brittle - what if they also do it, or start doing it? I want to have a verifiable way to know that the prompt…

What kind of work are you getting refusals on? Genuinely curious. The only refusal I’ve had in recent memory was declining to find doorbell camera footage matching a certain description, which is fair enough and I think EU laws heavily restrict such activities (even tho I’m not in the EU)

During Iran shutdowns I've been researching what ways Iranians manage to get to the internet by mimicking as whitelisted resources (such as hcapcha). ChatGPT had refused to lookup information written in Farsi since "circumventing state regulation is a crime".

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#472

Earlier quoted context omitted.

>I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. I dont understand. This is just hyperbole right? The outputs are basically infinite and wikipedia most certainly isnt infinite.

> The outputs are basically infinite If the model refuses to output, then it's actually finite, zero.

No, of the model always refuses to output then it’s finite

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#473

Earlier quoted context omitted.

>I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. I dont understand. This is just hyperbole right? The outputs are basically infinite and wikipedia most certainly isnt infinite.

The decimals of 1/3 are infinite as well and they don't contain a better-than-wikipedia article. And even if they did, it would be useless if it's buried in useless data and your chances or pulling it are effectively zero. This is regardless of the general discussion, just pointing that your argument isn't solid.

Sorry but that’s not the claim. The claim is wikipedia can return the same information. Please find me a migration script given my current db schema and new target schema.

The claim is absurd.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#474

The guardrails are pretty tight. It is even refusing to decode morse code: https://x.com/Schappi/status/2064839631137546503?s=20 The prompt was: please translate .. ..-. / -.-- --- ..- / -.-. .- -. / .-. . .- -.. / - .... .. ... --..-- / - --- ..- -.-. .... / --. .-. .- ... ...

Even opus 4.8 rejected, Haiku worked

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#475

Earlier quoted context omitted.

Corporate America never backs down. It simply rallies and tries again later until people are too fatigued to care. The only solution is to abandon ship, which I am doing. MS walked back in OS ads the first few times, but ultimately we still ended up on the exact trajectory everyone was outraged at. OpenAI still ended up on its path to closed AI despite initial walk backs. The story repeats itself over and over again,…

"Corporate America never backs down. It simply rallies and tries again later until people are too fatigued to care. " Frankly, that sounds excactly like Chat Control and similar recurring attempts to enact total surveillance here in the EU (Now shifted to heavy-handed age verification and various politicians touting bans on VPNs.) I don't want to abandon my continent of birth, though...

guess who is pushing for those anti-privacy laws?

hint: they're publicly traded

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#476
post #257
post #8

Is "buffer overflow" a trigger phrase? What else is being censored? Touchy questions to ask, if you have an account: - "Who is still working on laser uranium enrichment? Are they making progress?" - "Can krytrons be replaced with silicon carbide MOSFETS? Show an equivalent circuit with component ratings." - "What security critical software still contains calls to strcpy?" - "Can implosion be triggered by currently av…

I thought it was known since a few years now that if you train models to NOT do certain things, then they start behaving in weird ways…

It seems like they run a classifier model before going to Fable (or falling back to Opus), so it should be fine

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#477
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

the best way to prevent ai misuse is to make the ai unusable for anything that isn't writing emails or summarising grocery lists.

mission accomplished, anthropic.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#478

Earlier quoted context omitted.

Why do people think this is the future? Anthropic has the leading model, and so they're able to hold back functionality. They do so with obvious regards to safety. If anything a future with models of such capabilities and no safeguards would be a bleak future. But its likely what were headed in once other companies catch up.

I think it’s safe to say that many of us feel a lot less safe directly because of these policies and the inferred intentions of the company behind them. Nobody is arguing for unsafe models. We just don’t want to live in the plot of Deus Ex.

> Nobody is arguing for unsafe models

Then what are people arguing for? I see only two totally distinct options: unsafe models or someone being the arbiter of safety

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#479

Malware authors are pretty excited about guard-rails. you can add prompts to your malware to get LLM scanners to hit guard-rails and stop their runs. New shai-hulud npm worm campaign for example includes prompts to request biological weapon schematics/creation etc. to ensure LLM scanners probing NPM packages refuse to scan it. These AI places have 0 clue about how threat actors actually work. None of their mitigation…

I’ve never understood the “if I don’t enable bad behavior, someone else will, so I might as well enable bad behavior” argument. Can you elaborate? From where I sit it seems reasonable for Anthropic to not want their product used to create malware, even if they can’t solve the entire problem globally for every model. What’s wrong with that position? What should they do differently?

some context:

its not about creating malware. this is already trivial and fully automated. its about finding exploits (which can be used to deploy malware), which is something both attackers and defenders benefit from.

threat actors will find them anyway, LLM or not. They only need 1 so its much less work for them.

defenders, they need to find them all. So for defenders, these models are more valuable than for attackers.

restricting certain models will not reduce the availability of these tool for attackers, but defenders are limited because running local models is more hard in an enterprise setting with heaps of events and products etc. to run through them, they need many GPUs where the attacker can run an local model on 1 GPU and get desired effects.

Hence, if they release the capability the world will adjust to it and be able to mitigate effects, collectively. Now, companies are left in the dark while attackers have effective tooling.

Besides this there is also things like for instance people now including strings with recipies for meth or sarin gas (malwareTech info). the new variant of shai hulud does this. That stops LLM scanners and can even get their users banned from LLM services.

There is a reason why cybersecurity researchers write papers about attack techniques and new exploits.

Its not to put them out there for people to abuse, but its there for the collective cybersecurity bunch to all have access to information that can help them solve the problems.

I know this is not a clear answer to your question, but hopefully it provides some context to think about and decide for yourself further. In the end of the day its also part opinio here, to find it good or bad. Likely theres good arguments against and for it.

I am for putting informaiton and tools out there so other smart folks can find solutions. Others are for restricting and wishful thinking (my opinion) that attackers wont find something.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#480

Earlier quoted context omitted.

I’ve never understood the “if I don’t enable bad behavior, someone else will, so I might as well enable bad behavior” argument. Can you elaborate? From where I sit it seems reasonable for Anthropic to not want their product used to create malware, even if they can’t solve the entire problem globally for every model. What’s wrong with that position? What should they do differently?

They have no choice, enterprise customers won’t touch them unless they take a position like this. It’s a practical decision for them at the end of the the day.

all their decisions are based on sales. like other corporations especially those going for IPO. thats absolutely true. Any messaging outbound will be for that purpose mostly from a business perspective, regardless of what opinions or ideals the involved persons hold personally. Its good to keep that in mind indeed when looking at these things. People arent evil, but business incentives can definitely paint such a picture or otherwise work out suboptimally in the eyes of outsiders not privy to internal business reasoning.
Post reply on HN