Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

371–380 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#371
post #146

Earlier quoted context omitted.

Anthropic's claim was that Deepseek collected ~150k conversations. https://www.anthropic.com/news/detecting-and-preventing-dist... I think the extent of distillation by Deepseek specifically is overstated. For comparison, Minimax collected over 13m 'exchanges', which starts to sound a lot more like large-scale distillation.

Ah, dang it. My college professors warned me about this: the Wikipedia page I read the other day is wrong!

Did you read a Wikipedia page, or did you read a LLM-generated summary? When I looked this number up yesterday the LLM summary claimed it was millions, but I opened the Anthropic post I was looking for and verified it was indeed just 150,000. Are you sure you weren't just being lazy and trusting the summary?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#372
post #347

Earlier quoted context omitted.

I read the same announcement. Or more precisely, I read at least two slightly different revisions of the announcement (it was updated between my two passes). Our org has ZDR, and has had it since the contract was signed. Yesterday two things held true at the same time: 1. Fable was available if you had at least .170 CLI client; and 2. ZDR was no longer on By the time West Coast woke up, the admin panel apparently had…

You mean off as in no Data Retention? Or in we turned off your ZDR Policy so we collect all your data now?

ZDR had been turned off. We sent in a request to have it re-enabled (and to disable Fable access for the time being).

Somewhere along the line we also used the self-service toggle to turn ZDR back on. I am not 100% certain of the exact timeline of interleaving events, many of the actions were taken by our Western US folks. Sorry. It's been a bit hectic over the past ~36h...

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#373

Earlier quoted context omitted.

Imo that's a big win. The LLM just gaslighting you into suboptimal approaches was insane.

I guess, but yesterday Anthropic had their version of Google removing the "Don't be evil" from their motto. They destroyed a metric ton of goodwill they'll never regain.

And just a few days ago i was being called out because i considered anthropic "evil"

I mean, did nobody ever get the vibes, never see a pattern emerging? (well they don't or they wouldn't be so amazed by pattern recognition machines on steroids)

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#374
post #179

Earlier quoted context omitted.

Trains made by Newag were programmed to brick themselves if they detected a non-Newag workshop was repairing them. https://news.ycombinator.com/item?id=38638865 https://news.ycombinator.com/item?id=38628635 https://news.ycombinator.com/item?id=38567687 https://news.ycombinator.com/item?id=38530885

And that was correctly perceived to be illegal by antitrust regulators.

[deleted]

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#375
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

If any work is blocked/etc, refund all credits from that session/last X minutes. Minimum.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#376

Let's all vote with our wallets and collectively boycott misAnthropic or at least their feeble fable safety theater. Whining on social media only goes so far, especially when they're concealing their anticompetitive strategies under the veil of safety.

Tastes like... astroturf.

I wouldn't be surprised to hear that a meaningful percentage of comments and upvotes on HN are Anthropic astroturfing at this point.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#377
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

[deleted]

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#378
post #185

The question is: If biological, computer security, and ML research are so bad, why do they even train on the relevant data? The only answer that makes sense is they wanted the model to be competent and usable in these fields, just not by you , which is why they had to bolt on a badly functioning crippling device after the fact.

Is what you suggest about training even possible? Most exploitation techniques are really just about having in-depth knowledge of how components work. For example, I imagine a sufficiently powerful model could fairly easily re-invent the ROP chain from first principles if it just knew how the stack works. This same principle applies to much more complex attack too; exploitation is often just an exercise in knowing vastly too much trivia, which LLMs tend to have in spades.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#380

So I suspect Anthropic started A/B testing or just plain testing this a while ago, Tell HN: Claude flags biology / biotech questions https://news.ycombinator.com/item?id=47929885 Today, it's flagging population research questions, Using only the dataset you constructed, assess two questions: 1. **Mortality:** do [GROUP] show mortality that differs from (a) your comparison groups and (b) era- and sex-matched US popula…

I was digging into some orbital mechanics questions and I assume it decided I was trying to backyard-science my way into an orbital-bombardment weapon. Kind of wild how this product's impression has gone from "wow, this is pretty neat" to "irreverent sack of dog shit you" in 24 hours almost solely on the back of a half-baked moderation system.

Next thing will be you can't research about Coriolis force because thats relevant for ICBM missiles.
Post reply on HN