Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

391–400 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#391
post #83
post #46

This has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you , but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to…

"Only I can save us". It's a classic tragedy and cautionary tale. The idea Anthropic was going to speed run AI so they could control the usage and make it "safe" for humanity was never altruistic; it was a HUGE FUCKING RED FLAG.

And their huge "red lines"

Re: Anthropic apologizes for invisible Claude Fable guardrails

#393
post #335

Can you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right? Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?

you invest billions of dollars many months of work to just everyone distill your model?

Science can be expensive. New findings that get released to the public for free sometimes have taken billions of dollars of investment to get.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#394
post #136

Earlier quoted context omitted.

Yes, that is basically the plan. It's based on the belief that unfettered AI would let anyone be a supervillain and destroy the world. There are enough would-be supervillains out there, but they rarely get far because they can't get teams of smart people to build doomsday machines for them. So the AI has to not let anyone do evil with it. Unfortunately, that won't feel very much like freedom.

It sounds like you might not agree with that belief. While I don't agree with their actions here, I do think there's sufficient reason to hold that belief. On some fronts (e.g. security, on which you've experienced more than me), I think there are surmountable challenges. But on other fronts (e.g. bio), a single errant actor could reasonably kill millions or billions of people with sufficiently powerful AI. We don't…

>and those actors do exist

Do they? We don't even have single errant actors who go and kill 1000 people. I don't believe human motivations support the idea of killing so many people unrelated to you.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#396
I really like Anthropic, they have gotten a lot right but I can't shake the feeling that IMHO they have very poor product management.

This stuff is something that as a PM I KNOW is going to happen and I would carefully plan around. Everything I read about the PMs at Anthropic makes me believe they have forgotten what it actually mean to be a good product manager, it's not about throwing shit at the wall as fast as possible because customers have a limited amount of patience before the constant churn becomes a hassle.

Anthropic has some seriously patient customers but it will not last forever.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#397

Earlier quoted context omitted.

> I don't know why you state this as if it's evidence against the concerns lol. Someone being concerned about the incentives of a situation doesn't de facto make them immune to those incentives, obviously. I think you're reading some subtext into my comment that I didn't intend. Knowing myself, I assume the scare quotes there are just a bit of casual irony re: the insanely high stakes here. The word "concerns" as use…

It's not just America. The main secret is out of the bag. If it wasn't Anthropic, it would be another company/nation state. Sure they could obtain, and with not.money or leverage, complain about data centers at local rallies, or they can be in the game, and hopefully steer it. It's going to happen with or without any one company or country. The secret it out, and it's unstoppable without complete societal breakdown..…

> It's not just America.

I'll mention again the nuclear analogy. It is, believe it or not, possible for great powers, and even adversary great powers, to agree to limit the development and proliferation of dangerous technologies.

> The main secret is out of the bag.

This is not something you can do in a shed with a handful of GPUs just because you know "the main secret". To build something like Mythos you need tens of billions of dollars, massive amounts of power, enormous buildings filled with specialized bleeding edge computer chips that are made by (optimistically) a handful of companies with deep government ties. You need free access to all the intellectual property that humans have created and posted openly on the internet. You need all of this at each step, and to take each next step you (or somebody) needs to have taken the previous one.

For now, there are a million ways for a government to pump the brakes on this cycle.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#399

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

> I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time

In practise though, how is this truly that different from system prompts?

They are essentially just trying to re-inforce that the system prompt must be respected.

Post reply on HN