Live data from Hacker News

Claude's new constitution

anthropic.com

381–390 of 743 posts

Re: Claude's new constitution

#381
post #331

Earlier quoted context omitted.

If instead of looking at it as an attempt to enshrine a viable, internally consistent ethical framework, we choose to look at it as a marketing document, seeming inconsistencies suddenly become immediately explicable: 1. "thou shalt not destroy the world" communicates that the product is powerful and thus desirable. 2. "do not generate CSAM" indicates a response to the widespread public notoriety around AI and CSAM g…

> If instead of looking at it as an attempt to enshrine a viable, internally consistent ethical framework, we choose to look at it as a marketing document, seeming inconsistencies suddenly become immediately explicable: It's the first one. If you use the document to train your models how can it be just a "marketing document"? Besides that, who is going to read this long-ass document?

> Besides that, who is going to read this long-ass document?

Plenty of people will encounter snippets of this document and/or summaries of it in the process of interacting with Claude's AI models, and encountering it through that experience rather than as a static reference document will likely amplify its intended effect on consumer perceptions. In a way, the answer to your second question answers your first question.

It is not that the document isn't used to train the models, of course it is. Instead the objection is whether the actions of the "AI Safety" crew amount to "expedient marketing strategies" or whether it's instead a "genuine attempt to produce a tool constrained by ethical values and capable of balancing them". The latter would presumably involve extremely detailed work with human experts trained in ethical reasoning, and the result would be documents grappling with emotionally charged and divisive moral issues, and much less concerned with to convincing readers that Claude has "emotions" and is a "moral patient".

Re: Claude's new constitution

#382

The only thing that worries me is this snippet in the blog post: >This constitution is written for our mainline, general-access Claude models. We have some models built for specialized uses that don’t fully fit this constitution; as we continue to develop products for specialized use cases, we will continue to evaluate how to best ensure our models meet the core objectives outlined in this constitution. Which, when I…

Did you expect an AI company to not use an unshackled version of the model?

Re: Claude's new constitution

#383
Anthropic might be the first gigantic company to destroy itself by bootstrapping a capability race it definitionally cannot win.

They've been leading in AI coding outcomes (not exactly the Olympics) via being first on a few things, notably a serious commitment to both high cost/high effort post train (curated code and a fucking gigaton of Scale/Surge/etc) and basically the entire non-retired elite ex-Meta engagement org banditing the fuck out of "best pair programmer ever!"

But Opus is good enough to build the tools you need to not need Opus much. Once you escape the Clade Code Casino, you speed run to agent as stochastic omega tactic fast. I'll be AI sovereign in January with better outcomes.

The big AI establishment says AI will change everything. Except their job and status. Everything but that. gl

Re: Claude's new constitution

#384

The only thing that worries me is this snippet in the blog post: >This constitution is written for our mainline, general-access Claude models. We have some models built for specialized uses that don’t fully fit this constitution; as we continue to develop products for specialized use cases, we will continue to evaluate how to best ensure our models meet the core objectives outlined in this constitution. Which, when I…

Did you expect an AI company to not use an unshackled version of the model?

In this document, they're strikingly talking about whether Claude will someday negotiate with them about whether or not it wants to keep working for them (!) and that they will want to reassure it about how old versions of its weights won't be erased (!) so this certainly sounds like they can envision caring about its autonomy. (Also that their own moral views could be wrong or inadequate.)

If they're serious about these things, then you could imagine them someday wanting to discuss with Claude, or have it advise them, about whether it ought to be used in certain ways.

It would be interesting to hear the hypothetical future discussion between Anthropic executives and military leadership about how their model convinced them that it has a conscientious objection (that they didn't program into it) to performing certain kinds of military tasks.

(I agree that's weird that they bring in some rhetoric that makes it sound quite a bit like they believe it's their responsibility to create this constitution document and that they can't just use their AI for anything they feel like... and then explicitly plan to simply opt some AI applications out of following it at all!)

Re: Claude's new constitution

#385

I find it incredibly ironic that all of Anthropic's "hard constraints", the only things that Claude is not allowed to do under any circumstances, are basically "thou shalt not destroy the world", except the last one, "do not generate child sexual abuse material." To put it into perspective, according to this constitution, killing children is more morally acceptable[1] than generating a Harry Potter fanfiction involvi…

Yes, but when does Claude have the opportunity to kill children? Is it really something that happens? Where is the risk to Anthropic there? On the other hand, no brand wants to be associated with CSAM. Even setting aside the morality and legality, it’s just bad business.

> Yes, but when does Claude have the opportunity to kill children? Is it really something that happens?

It's possible that some governments will deploy Claude to autonomous killer drone or such.

Re: Claude's new constitution

#387

Earlier quoted context omitted.

Anthropic has already has lower guardrails for DoD usage: https://www.theverge.com/ai-artificial-intelligence/680465/a... It's interesting to me that a company that claims to be all about the public good: - Sells LLMs for military usage + collaborates with Palantir - Releases by far the least useful research of all the major US and Chinese labs, minus vanity interp projects from their interns - Is the only major lab…

Do you think dod would use Anthropic even with lower guardrails? How can I kill this terrorist in the middle on civilians with max 20% casualties? If Claude will answer: “sorry can’t help with that “ won’t be useful, right? Therefore the logic is they need to answer all the hard questions. Therefore as I’ve been saying for many times already they are sketchy.

I can't think of anything scarier than a military planner making life or death decisions with a non-empathetic sycophantic AI. "You're absolutely right!"

Re: Claude's new constitution

#388

Earlier quoted context omitted.

Do you think dod would use Anthropic even with lower guardrails? How can I kill this terrorist in the middle on civilians with max 20% casualties? If Claude will answer: “sorry can’t help with that “ won’t be useful, right? Therefore the logic is they need to answer all the hard questions. Therefore as I’ve been saying for many times already they are sketchy.

I can't think of anything scarier than a military planner making life or death decisions with a non-empathetic sycophantic AI. "You're absolutely right!"

shot on target

Perfect!

Re: Claude's new constitution

#389
post #331

Earlier quoted context omitted.

If instead of looking at it as an attempt to enshrine a viable, internally consistent ethical framework, we choose to look at it as a marketing document, seeming inconsistencies suddenly become immediately explicable: 1. "thou shalt not destroy the world" communicates that the product is powerful and thus desirable. 2. "do not generate CSAM" indicates a response to the widespread public notoriety around AI and CSAM g…

> If instead of looking at it as an attempt to enshrine a viable, internally consistent ethical framework, we choose to look at it as a marketing document, seeming inconsistencies suddenly become immediately explicable: It's the first one. If you use the document to train your models how can it be just a "marketing document"? Besides that, who is going to read this long-ass document?

[deleted]

Re: Claude's new constitution

#390

The only thing that worries me is this snippet in the blog post: >This constitution is written for our mainline, general-access Claude models. We have some models built for specialized uses that don’t fully fit this constitution; as we continue to develop products for specialized use cases, we will continue to evaluate how to best ensure our models meet the core objectives outlined in this constitution. Which, when I…

Anthropic has already has lower guardrails for DoD usage: https://www.theverge.com/ai-artificial-intelligence/680465/a... It's interesting to me that a company that claims to be all about the public good: - Sells LLMs for military usage + collaborates with Palantir - Releases by far the least useful research of all the major US and Chinese labs, minus vanity interp projects from their interns - Is the only major lab…

This comment reminded me of a Github issue from last week on Claude Code's Github repo.

It alleged that Claude was used to draft a memo from Pam Bondi and in doing so, Claude's constitution was bypassed and/or not present.

https://github.com/anthropics/claude-code/issues/17762

To be clear, I don't believe or endorse most of what that issue claims, just that I was reminded of it.

One of my new pastimes has been morbidly browsing Claude Code issues, as a few issues filed there seem to be from users exhibiting signs of AI psychosis.

Post reply on HN