This has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you , but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to…
"Only I can save us". It's a classic tragedy and cautionary tale. The idea Anthropic was going to speed run AI so they could control the usage and make it "safe" for humanity was never altruistic; it was a HUGE FUCKING RED FLAG.
Anthropic apologizes for invisible Claude Fable guardrails
391–400 of 489 posts
Re: Anthropic apologizes for invisible Claude Fable guardrails
#392Re: Anthropic apologizes for invisible Claude Fable guardrails
#393Can you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right? Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?
you invest billions of dollars many months of work to just everyone distill your model?
Re: Anthropic apologizes for invisible Claude Fable guardrails
#394Earlier quoted context omitted.
Yes, that is basically the plan. It's based on the belief that unfettered AI would let anyone be a supervillain and destroy the world. There are enough would-be supervillains out there, but they rarely get far because they can't get teams of smart people to build doomsday machines for them. So the AI has to not let anyone do evil with it. Unfortunately, that won't feel very much like freedom.
It sounds like you might not agree with that belief. While I don't agree with their actions here, I do think there's sufficient reason to hold that belief. On some fronts (e.g. security, on which you've experienced more than me), I think there are surmountable challenges. But on other fronts (e.g. bio), a single errant actor could reasonably kill millions or billions of people with sufficiently powerful AI. We don't…
Do they? We don't even have single errant actors who go and kill 1000 people. I don't believe human motivations support the idea of killing so many people unrelated to you.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#395Re: Anthropic apologizes for invisible Claude Fable guardrails
#396This stuff is something that as a PM I KNOW is going to happen and I would carefully plan around. Everything I read about the PMs at Anthropic makes me believe they have forgotten what it actually mean to be a good product manager, it's not about throwing shit at the wall as fast as possible because customers have a limited amount of patience before the constant churn becomes a hassle.
Anthropic has some seriously patient customers but it will not last forever.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#397Earlier quoted context omitted.
> I don't know why you state this as if it's evidence against the concerns lol. Someone being concerned about the incentives of a situation doesn't de facto make them immune to those incentives, obviously. I think you're reading some subtext into my comment that I didn't intend. Knowing myself, I assume the scare quotes there are just a bit of casual irony re: the insanely high stakes here. The word "concerns" as use…
It's not just America. The main secret is out of the bag. If it wasn't Anthropic, it would be another company/nation state. Sure they could obtain, and with not.money or leverage, complain about data centers at local rallies, or they can be in the game, and hopefully steer it. It's going to happen with or without any one company or country. The secret it out, and it's unstoppable without complete societal breakdown..…
I'll mention again the nuclear analogy. It is, believe it or not, possible for great powers, and even adversary great powers, to agree to limit the development and proliferation of dangerous technologies.
> The main secret is out of the bag.
This is not something you can do in a shed with a handful of GPUs just because you know "the main secret". To build something like Mythos you need tens of billions of dollars, massive amounts of power, enormous buildings filled with specialized bleeding edge computer chips that are made by (optimistically) a handful of companies with deep government ties. You need free access to all the intellectual property that humans have created and posted openly on the internet. You need all of this at each step, and to take each next step you (or somebody) needs to have taken the previous one.
For now, there are a million ways for a government to pump the brakes on this cycle.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#398I moved off Claude Code 3 months ago. That decision keeps getting better and better as time goes on.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#399I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…
In practise though, how is this truly that different from system prompts?
They are essentially just trying to re-inforce that the system prompt must be respected.