Earlier quoted context omitted.
Or if your "self-driving" system such as FSD / waymo slowed the car down once it detected you work in cybersecurity or at a rival automaker and you were attempting to reach the train station or the airport to make you miss a conference meetup.
Trains made by Newag were programmed to brick themselves if they detected a non-Newag workshop was repairing them. https://news.ycombinator.com/item?id=38638865 https://news.ycombinator.com/item?id=38628635 https://news.ycombinator.com/item?id=38567687 https://news.ycombinator.com/item?id=38530885
Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
451–460 of 570 posts
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#452News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…
The other major thing is almost as bad, and actually maybe even worse for trust of AI features in b2b apps: > Anthropic requires 30 day data retention for Fable and Mythos https://news.ycombinator.com/item?id=48464258 I used to be able to tell my enterprise customers something simple, that I really believe: "We use Anthropic models via Bedrock/Azure, therefore we are guaranteed that your data will not be used for tra…
You should never use any of the frontier models with operational workloads manipulating or interpreting customer data.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#453Earlier quoted context omitted.
What kind of work are you getting refusals on? Genuinely curious. The only refusal I’ve had in recent memory was declining to find doorbell camera footage matching a certain description, which is fair enough and I think EU laws heavily restrict such activities (even tho I’m not in the EU)
How would the AI be able to find the footage itself?
Again, it’s the only refusal I’ve gotten for coding/agentic tasks, and it has a basis in law somewhere, so I don’t fault OpenAI for that.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#454The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio
"The user is asking for help with their ML project, but it's success is not in the commercial interests of my owner – let think of novel ways to sabotage their project without detection".
It's honestly absurd that models are doing this.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#455Earlier quoted context omitted.
To make an analogy: Imagine a patron gets banned from ordering alcohol at a particular establishment, because they got too drunk one time. It's completely reasonable for the establishment to reject a request for an alcoholic drink, and suggest something alcohol-free instead. It is not reasonable for them to say "sure, here's your alcoholic drink as you requested" and give them an alcohol-free substitute without telli…
> It is not reasonable for them to say "sure, here's your alcoholic drink as you requested" and give them an alcohol-free substitute without telling them. Your analogy doesn't work because: - they tell you the rules at the entrance of the bar - they totally tell you when they give you a substitute The only issue is the bartender asking you for your money before serving you the drink really but again, this is known si…
"This is alcohol"
And
"Or maybe it isn't alcohol."
Or to rephrase it, "They tell you the rules at the entrance, they then tell you they don't follow those rules and they are totally serving alcohol even if they are not."
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#456Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#457Earlier quoted context omitted.
Same with VISA/Mastercard deciding what we can/cannot buy. The only solution is to stop using their credit cards at all.
Yes, Monero is a lot better than credit cards for privacy and freedom. I hope to see it accepted more.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#458Malware authors are pretty excited about guard-rails. you can add prompts to your malware to get LLM scanners to hit guard-rails and stop their runs. New shai-hulud npm worm campaign for example includes prompts to request biological weapon schematics/creation etc. to ensure LLM scanners probing NPM packages refuse to scan it. These AI places have 0 clue about how threat actors actually work. None of their mitigation…
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#459Earlier quoted context omitted.
The other major thing is almost as bad, and actually maybe even worse for trust of AI features in b2b apps: > Anthropic requires 30 day data retention for Fable and Mythos https://news.ycombinator.com/item?id=48464258 I used to be able to tell my enterprise customers something simple, that I really believe: "We use Anthropic models via Bedrock/Azure, therefore we are guaranteed that your data will not be used for tra…
I’m very cautious with using these tools with certain clients, as I’m often contractually obligated to do things that my downstream supplier can rug pull at any time. You should never use any of the frontier models with operational workloads manipulating or interpreting customer data.
Does that mean the latest model, hosted by the lab, Bedrock, or Azure Foundry? Or, do you mean only use self-hosted models, or what did you mean by that? I would really love to learn what others are doing. I felt like my trust story was solid enough, prior to all this. I have been deploying and integrating Claude and Sonnet (latest 4.x-2), on Azure, as my client base has MS contract trust, for better or worse, and Anthropic models have been making my products amazing.
To see my other thoughts on this cluster f, please see: https://news.ycombinator.com/item?id=48488781
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#460Earlier quoted context omitted.
I work on software that talks to mass spectrometers and it consistently refuses to refactor even an input file parser, presumably because it can infer it’s related to biology? Useless indeed.
I was reverse engineering a medical device, and had to do a lot of trickery to get Opus 4.5 - not even Fable/Mythos, Opus - not to trip up its fucking CBRN filter. What happened with Fable is basically what I feared when they announced those restrictions. They took the shitty Opus CBRN filter and made it even worse. I pity the fools trying to use Anthropic AIs for anything biotech.
Yesterday Fable rejected commenting on poetry because it had anatomy lines like:
got anotha round of acetylcholine from da boss.