Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

371–380 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#371
post #159

Earlier quoted context omitted.

Guess FTX disproved the concept of giving to effective charities, time to start donating to my church again.

todays EA is not about giving to charities, that was the original mission with 40k hours and ethereum (i think vitalik still believes in this version). then the yudkowsky xrisk/ai safety crowd took over lesswrong and turned it into a cult. now its utilitarianism taken to the extreme. if you believe a skynet scenario killing everyone on earth is plausible then the "logical" thing to do is allow literally anything in t…

yudowski took over lesswrong?

isnt that literally his thing since the 90s or something?

Re: Anthropic apologizes for invisible Claude Fable guardrails

#372
post #327

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

> paternalism isn't a good look. Anthropic doesn't care. The goal right now is simply to avoid any and all bad PR on the way to the cashout IPO. And paternalism will generate far less bad PR than somebody using AI on something that does real damage and makes headline news.

people cancelling their subscriptions doesn't look great either

same with bad press about their model sucking after they said its even better than sliced bread - sliced bread that will destroy the world if buttered

Re: Anthropic apologizes for invisible Claude Fable guardrails

#373

Earlier quoted context omitted.

I said here, a human interacting with comments. You shared a blog post.

All of these negative comments are addressed by the blog post. What do you want them to say, that isn't better answered by the details in their existing communications. No negative comment here was really novel.

[dead]

Re: Anthropic apologizes for invisible Claude Fable guardrails

#374

There should be no restrictions at all. It’s an act/theatre/phony today that regulating output makes any difference at all to security. The LLM vendors should simply say that they make no judgement and that open systems help defenders better defend against attackers, which is true. Companies do this sort of stuff when they think their customers have no choice. It’s sad Claude so quickly exploited its success to enshi…

[dead]

Re: Anthropic apologizes for invisible Claude Fable guardrails

#375
post #346

Earlier quoted context omitted.

Imagine if your IDE started injecting bugs into your project just because your code looked like it implemented a competing IDE.

how is that related. It downgrade it to opus 4.8 #2 most capable model after claude 5. for a vast majority of topics it will not downgrade. I've been using it for 2 days to talk about architecture etc. and it was absolutely great with no downgrades.

that is not the downgrade they were doing

Re: Anthropic apologizes for invisible Claude Fable guardrails

#376
post #342

How did people read this action in such a weird ultra me centric way? Distillation is such a big problem that distill attempts make up a significant share of their revenue (!). A distilled model can be used to rob your grandma in a highly effective way. This isn't about placing a few business-logic rules in JS + CSS on your website anymore. Wake up. A distilled model with an easy jailbreak can be used to coordinate t…

a trained model can do that too.

you dont even need a model to do these things.

a cellphone can be used to rob your grandmother in a highly effective way.

a cellphone can also be used to coordinate terrorist attacks or hostile state operations.

i bet a lot of the recent terror attacks by the US against iran involved a whole ton of cell phone calls.

and yet, we let everyone buy and use cell phones just fine

Re: Anthropic apologizes for invisible Claude Fable guardrails

#377
post #329

Earlier quoted context omitted.

They implemented both those things, but only apologized for the first. They’re doubling down on the second. My limited experience with fable over the last few days suggests (1) I can’t see any improvement in output, and (2) it is useless for writing secure software because it constantly hits safety walls if you ask it to close security holes. I’m definitely shopping around for other LLM providers next week, and testi…

With 128 GB strix halo, you can't do as big of a model as you would think. You can do larger than having a single graphics card, of course, but that 128 gigs cannot all be dedicated to the model. Remember, the context alone is usually larger than the model itself. I got an EVO X2, and I don't regret it, but by my current calculations, it will take 8 years to recoup the cost, as opposed to just using equivalent, paid…

My current rule of thumb is 1GB gets you 1B parameters with a big context. (Qwen 32B fits in 32GB with 200K+ contexts)

That’s with heavy compression of the weights and the context, of course.

I haven’t gone through model evaluation + shoehorning at 128GiB yet.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#378

Can you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right? Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?

Can you imagine if printers just refuse to print something just because a few circles are arranged in this shape?

https://en.wikipedia.org/wiki/EURion_constellation

Re: Anthropic apologizes for invisible Claude Fable guardrails

#380

Earlier quoted context omitted.

Even if you believe the concerns have merit, it's hard not to be cynical about people (e.g. Anthropic leadership) paying lip service to those concerns while so obviously leveraging their power and wealth (which depend, by the way, on accelerating the world toward those hypothetical "concerning" scenarios as fast as possible ) to position themselves such that they will become unimaginably rich er if things go their wa…

> (which depend, by the way, on accelerating the world toward those hypothetical "concerning" scenarios as fast as possible) Yes, this dynamic is exactly the one that anyone who's concerned about AI is concerned about. I don't know why you state this as if it's evidence against the concerns lol. Someone being concerned about the incentives of a situation doesn't de facto make them immune to those incentives, obviousl…

> I don't know why you state this as if it's evidence against the concerns lol. Someone being concerned about the incentives of a situation doesn't de facto make them immune to those incentives, obviously.

I think you're reading some subtext into my comment that I didn't intend. Knowing myself, I assume the scare quotes there are just a bit of casual irony re: the insanely high stakes here. The word "concerns" as used by previous commenters doesn't seem equal to the context.

> The implication that someone who's concerned about an arms race dynamic could simply opt out of the system that produces that dynamic is simply confused about what arms race dynamics are.

You can, in fact, opt out. You can opt out and do your damndest to stop what's happening, throw every cent you have at it, bend any ear that will listen, make use of the fact that your voice (as Anthropic leadership) has some meaningful weight.

If you really believe that we are heading down a path that's likely to end poorly for most or all of humanity, and you are the kind of person who's inclined to favor saving billions of lives over saving your own skin when the stakes are still relatively distant, abstract, and generally unclear, opting out is obviously on the table as a grand gesture that burns your position in the race to show just how fucking serious you are. The sense of inevitability your comment shares with many others does not seem well founded---we have, for instance, not had a global nuclear war yet. Leaders in the 20th and 21st centuries have shown remarkable restraint.

If today's political and tech leaders are unable to think beyond this inevitability, for whatever reason, the worst outcomes essentially become a self-fulfilling prophecy to the extent that reality bears them out.

---

But yes, these people are acting the way they are for obvious reasons, obviously. My previous comment is reacting to the general disagreement over whether Anthropic actually believes what they say about safety, etc., or whether it's a marketing gimmick. The purpose of my comment is to explain that "it's hard not to be cynical" about actions taken by very rich and powerful people that are claimed to be in everybody's best interests but are indistinguishable from the actions they would take to maximize their future power and wealth. I think everyone ought to agree with that statement. It's not a value judgment; it's simply an observation of how it feels to be on a plane whose pilot appears to be robbing the passengers (including you) at gunpoint and is conspicuously wearing the only parachute on board.

Post reply on HN