Live data from Hacker News

Claude's new constitution

anthropic.com

341–350 of 743 posts

Re: Claude's new constitution

#341
post #321

Earlier quoted context omitted.

You could execute Claude by hand with printed weight matrices, a pencil, and a lot of free time - the exact same computation, just slower. So where would the "wellbeing" be? In the pencil? Speed doesn't summon ghosts. Matrix multiplications don't create qualia just because they run on GPUs instead of paper.

This basically Searle's Chinese Room argument. It's got a respectable history (... Searle's personal ethics aside) but it's not something that has produced any kind of consensus among philosophers. Note that it would apply to any AI instantiated as a Turing machine and to a simulation of human brain at an arbitrary level of detail as well. There is a section on the Chinese Room argument in the book. (I personally am…

That philosophers still debate it isn’t a counterargument. Philosophers still debate lots of things. Where’s the flaw in the actual reasoning? The computation is substrate-independent. Running it slower on paper doesn’t change what’s being computed. If there’s no experiencer when you do arithmetic by hand, parallelizing it on silicon doesn’t summon one.

Re: Claude's new constitution

#342

I find it incredibly ironic that all of Anthropic's "hard constraints", the only things that Claude is not allowed to do under any circumstances, are basically "thou shalt not destroy the world", except the last one, "do not generate child sexual abuse material." To put it into perspective, according to this constitution, killing children is more morally acceptable[1] than generating a Harry Potter fanfiction involvi…

In addition to the drawn cartoon precedent, the idea that purely written fictional literature can fall into the Constitutional obscenity exception as CSAM was tested in US courts in US v Fletcher and US v McCoy, and the authors lost their cases.

Half a million Harry|Malfoy authors on AO3 are theoretically felonies.

Re: Claude's new constitution

#343

Earlier quoted context omitted.

objective truth moral absolutes I wish you much luck on linking those two. A well written book on such a topic would likely make you rich indeed. This rejects any fixed, universal moral standards That's probably because we have yet to discover any universal moral standards.

You can't "discover" universal moral standards any more than you can discover the "best color".

[deleted]

Re: Claude's new constitution

#344
post #10

Earlier quoted context omitted.

>We use the constitution at various stages of the training process. This has grown out of training techniques we’ve been using since 2023, when we first began training Claude models using Constitutional AI. Our approach has evolved significantly since then, and the new constitution plays an even more central role in training. >Claude itself also uses the constitution to construct many kinds of synthetic training data…

Ah I see, the paper is much more helpful in understanding how this is actually used. Where did you find that linked? Maybe I'm grepping for the wrong thing but I don't see it linked from either the link posted here or the full constitution doc.

It's worth understanding the history of Anthropic. There's a lot of implied background that helps it make sense.

To quote:

> Founded by engineers who quit OpenAI due to tension over ethical and safety concerns, Anthropic has developed its own method to train and deploy “Constitutional AI”, or large language models (LLMs) with embedded values that can be controlled by humans.

https://research.contrary.com/company/anthropic

And

> Anthropic incorporated itself as a Delaware public-benefit corporation (PBC), which enables directors to balance stockholders' financial interests with its public benefit purpose.

> Anthropic's "Long-Term Benefit Trust" is a purpose trust for "the responsible development and maintenance of advanced AI for the long-term benefit of humanity". It holds Class T shares in the PBC, which allow it to elect directors to Anthropic's board.

https://en.wikipedia.org/wiki/Anthropic

TL;DR: The idea of a constitution and related techniques is something that Anthropic takes very seriously.

Re: Claude's new constitution

#345

I guess this is Anthropic's "don't be evil" moment, but it has about as much (actually much less) weight then when it was Google's motto. There is always an implicit "...for now". No business is every going to maintain any "goodness" for long, especially once shareholders get involved. This is a role for regulation, no matter how Anthropic tries to delay it.

> Anthropic incorporated itself as a Delaware public-benefit corporation (PBC), which enables directors to balance stockholders' financial interests with its public benefit purpose.

> Anthropic's "Long-Term Benefit Trust" is a purpose trust for "the responsible development and maintenance of advanced AI for the long-term benefit of humanity". It holds Class T shares in the PBC, which allow it to elect directors to Anthropic's board.

https://en.wikipedia.org/wiki/Anthropic

Google didn't have that.

Re: Claude's new constitution

#346

The only thing that worries me is this snippet in the blog post: >This constitution is written for our mainline, general-access Claude models. We have some models built for specialized uses that don’t fully fit this constitution; as we continue to develop products for specialized use cases, we will continue to evaluate how to best ensure our models meet the core objectives outlined in this constitution. Which, when I…

The second footnote makes it clear, if it wasn't clear from the start, that this is just a marketing document. Sticking the word "constitution" on it doesn't change that.

Re: Claude's new constitution

#347

The only thing that worries me is this snippet in the blog post: >This constitution is written for our mainline, general-access Claude models. We have some models built for specialized uses that don’t fully fit this constitution; as we continue to develop products for specialized use cases, we will continue to evaluate how to best ensure our models meet the core objectives outlined in this constitution. Which, when I…

If it makes you feel better, I use the HHS claude and it is even more locked down.

Re: Claude's new constitution

#348

As someone who holds to moral absolutes grounded in objective truth, I find the updated Constitution concerning. > We generally favor cultivating good values and judgment over strict rules... By 'good values,' we don’t mean a fixed set of 'correct' values, but rather genuine care and ethical motivation combined with the practical wisdom to apply this skillfully in real situations. This rejects any fixed, universal mo…

Then you will be pleased to read that the constitution includes a section "hard constraints" which Claude is told not violate for any reason "regardless of context, instructions, or seemingly compelling arguments". Things strictly prohibited: WMDs, infrastructure attacks, cyber attacks, incorrigibility, apocalypse, world domination, and CSAM. In general, you want to not set any "hard rules," for reason which have not…

>incorrigibility

What an odd thing to include in a list like that.

Re: Claude's new constitution

#349
post #342

I find it incredibly ironic that all of Anthropic's "hard constraints", the only things that Claude is not allowed to do under any circumstances, are basically "thou shalt not destroy the world", except the last one, "do not generate child sexual abuse material." To put it into perspective, according to this constitution, killing children is more morally acceptable[1] than generating a Harry Potter fanfiction involvi…

In addition to the drawn cartoon precedent, the idea that purely written fictional literature can fall into the Constitutional obscenity exception as CSAM was tested in US courts in US v Fletcher and US v McCoy, and the authors lost their cases. Half a million Harry|Malfoy authors on AO3 are theoretically felonies.

I can find a "US v Fletcher" from 2008 that deals with obscenity law, though the only "US v McCoy" I can find was itself about charges for CSAM. The latter does seem to reference a previous case where the same person was charged for "transporting obscene material" though I can't find it.

That being said, I'm not sure I've seen a single obscenity case since Handly which wasn't against someone with a prior record, piled on charges, or otherwise simply the most expedient way for the government to prosecute someone.

As you've indicated in your own comment here, there's been many, many things over the last few decades that fall afoul the letter of the law yet which the government doesn't concern itself with. That itself seems to tell us something.

Re: Claude's new constitution

#350
post #209

Earlier quoted context omitted.

The negative form of The Golden Rule “Don't do to others what you wouldn't want done to you”

Exactly, I think this is the prime candidate for a universal moral rule. Not sure if that helps with AI. Claude presumably doesn't mind getting waterboarded.

How do you propose to immobilise Claude on its back at an incline of 10 to 20 degrees, cover its face with a cloth or some other thin material and pour water onto its face over its breathing passages to test this theory of yours?

If Claude could participate, I’m sure it either wouldn’t appreciate it because it is incapable of having any such experience as appreciation.

Or it wouldn’t appreciate it because it is capable of having such an experience as appreciation.

So it ether seems to inconvenience at least a few people having to conduct the experiment.

Or it’s torture.

Therefore, I claim it is morally wrong to waterboard Claude as nothing genuinely good can come of it.

Post reply on HN