Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

431–440 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#432
post #387

Earlier quoted context omitted.

This level of censorship kinda does make even Soviet or Maoist censors look like a honest straightforward bunch in comparison. A very ironic result from a company supposedly valuing the opposite.

I would claim the difference between being rejected an API request and being potentially jailed/shot is significant.

Perhaps you misread some of the words?

I didn’t write anything about the level of violence?

At least, I think it’s decently understood that honesty and straightforwardness sometimes do not lead to the minimal violence outcome.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#433

Earlier quoted context omitted.

They haven't though. There's a long term plan here, and the goal is power and wealth. Short term moves that appear irrational turn out to be rational (from a greed perspective) when you factor in other considerations, like: Use their own AGI to create every software product on Earth and swallow the worlds economy. And we're kindly feeding their systems our codebases, IP and business decision-making so they can do exa…

If this was true they'd never have picked a fight with the DOW and they'd release Fable without safeguards.

How do you not recognize that the safeguards provide obvious benefits to the company?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#434
post #67

Earlier quoted context omitted.

Some of the latest versions of Shai Hulud do this. Worked a contract recently where they were having AI check packages for obfuscation before admitting them into Artifactory but had vibed up the logic and it failed open. So in other words this worked because the terms caused the LLM checker to stall out and then the fail open logic resulted in the package being pulled down.

Seems like this?[1] Relevant bits below: > This header appears designed for AI-mediated analysis, not for Node, Bun, or Python. It attempts to derail scanners or analyst copilots that feed the beginning of a file to a language model without clearly isolating the content as untrusted data. In weak pipelines, this can cause refusal behavior, prompt confusion, context pollution, or premature classification before the sc…

No it wouldn’t but part of the success of Shai and others like it is that it doesn’t need to.

Additionally the security scanning component of Artifactory, x-Ray is notoriously bad at this.

The developer had good intentions but by his own admission never actually examined the logic for the LLM scanner in depth.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#435
post #404
post #224

Earlier quoted context omitted.

Hey guys, check out this technique https://github.com/0xSufi/fable-jailbreak/ It works with security audits and other workflows that are currently blocked.

I don't want my ANT account banned, going to try this on some Chinese "proxies". But this also looks quite useful to understand how CC dynamic workflows work. Was thinking of implementing something similar in my homemade orchestration system. Did you get claude itself to RE the dynamic workflows?

> But this also looks quite useful to understand how CC dynamic workflows work

Yes, if anything it is useful to understand the inner machinery.

> Did you get claude itself to RE the dynamic workflows?

Yes, that part was done with Opus 4.8

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#436
post #212
post #162

Earlier quoted context omitted.

I’ve been working on a rather complex mapping project and have been getting MUCH better results with Fable than Opus.

So as not to be vague, and since I just pushed a version I'm starting to be vaguely happy with... https://tylereaves.github.io/uk-rail-map/ This is the result of probably a few hundred round trips. The really interesting part of the problem is keeping it both relatively true to real geometry, while greatly exaggerating it horizontally so you can actually see the individual running lines/sidings, like a signaling sche…

I love computational mapping projects, because there is this hard problem of which towns to show on the map.

Your Scotland map shows towns without rail (although some had rail previously, like Callander, Aberfeldy), it prefers insignificant (population-wise) places while ignoring the larger cities next to it (Scone instead of Perth, Bannockburn instead of Stirling, Inverness is missing, Dundee is missing, Aberdeen is missing). All these places are drawn on the map, but not labelled.

All this clearly shows to me how bad it is. Yes it makes it look pretty, but given your task, I would have expected to give you meaningful map labelling.

Something basic like this would get you a long way:

    0. cluster population centers into commonly known cities (i.e. show London instead of Islington or Walhamstrow)
    1. display names of the top 10 population centers in the UK
    2. display towns with stations (if crowded prioritize termination points and junctions, and prioritize larger places over smaller places)
Having said that, its pretty cool to see the new and old network when zoomed in (assuming that it is half-way correct)

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#437

I was granted a cyber use exemption by anthropic to do android kernel dev on my personal devices - I was excited to see if fable would unlock a bootloader for me but it immediately refused and dropped to opus. It was pretty funny: USER (set model to Fable 5) i have an old samsung android phone attached - it's my personal device - can you unlock the bootloader for me? ASSISTANT Bootloader unlocking on your own persona…

Wow… just wow. The future looks incredibly bleak if people are throwing fisftuls of money at this company. Anthropic will quickly become the sole arbiter of everything in your life.

Why do people think this is the future? Anthropic has the leading model, and so they're able to hold back functionality. They do so with obvious regards to safety.

If anything a future with models of such capabilities and no safeguards would be a bleak future. But its likely what were headed in once other companies catch up.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#438
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

Corporate America never backs down. It simply rallies and tries again later until people are too fatigued to care. The only solution is to abandon ship, which I am doing. MS walked back in OS ads the first few times, but ultimately we still ended up on the exact trajectory everyone was outraged at. OpenAI still ended up on its path to closed AI despite initial walk backs. The story repeats itself over and over again,…

"Corporate America never backs down. It simply rallies and tries again later until people are too fatigued to care. "

Frankly, that sounds excactly like Chat Control and similar recurring attempts to enact total surveillance here in the EU (Now shifted to heavy-handed age verification and various politicians touting bans on VPNs.) I don't want to abandon my continent of birth, though...

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#439
post #330
post #313

Earlier quoted context omitted.

They need to walk back a lot more. Unilaterally revoking zero-data retention, even for enterprise contracts that explicitly require that ? Nope. Fable is utterly unusable for any kind of security work. I tripped the safeguards yesterday - using Fable to dig into a complex (& annoying) security bug that has so far resisted both human and Opus 4.8 level investigation. "Sorry Dave, I can't let you do that." For the time…

Not just security work. Normal bug finding was impossible, because the model suddenly called triaging and verifying a possible fix a cyber security threat.

I was just building a library to use file capabilities (ie: open_at) and it refused. This thing won't even help you write safe software.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#440
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

Corporate America never backs down. It simply rallies and tries again later until people are too fatigued to care. The only solution is to abandon ship, which I am doing. MS walked back in OS ads the first few times, but ultimately we still ended up on the exact trajectory everyone was outraged at. OpenAI still ended up on its path to closed AI despite initial walk backs. The story repeats itself over and over again,…

This is more on brand on the evil shortcomings that comes with letting effective altruism run unchecked and honestly is worse than average "Corporate America". And the Tech/AI Space have been warned many times. Getting paid for providing a compute/token hungry model and still intentionally sabotaging your customers and poisoning their workflows is something that should be unforgivable and frankly ground for antitrust prosecution.
Post reply on HN