Live data from Hacker News

Claude's new constitution

anthropic.com

601–610 of 743 posts

Re: Claude's new constitution

#602

The constitution contains 43 instances of the word 'genuine', which is my current favourite marker for telling if text has been written by Claude. To me it seems like Claude has a really hard time _not_ using the g word in any lengthy conversation even if you do all the usual tricks in the prompt - ruling, recommending, threatening, bribing. Claude Code doesn't seem to have the same problem, so I assume the system pr…

This is a great (and funny) thread but for anyone too lazy to read the actual constitution and still curious about this, they directly state that Claude wrote first drafts for several of the human authors of the document.

Re: Claude's new constitution

#603
post #586

* Anthropic accepted a 200M contract from the US Department of Defence * Anthropic seeked contracts from the United Arab Emirates and Qatar, the leaked memo acknowledges that the contracts will enrich dictators * Anthropic spent more than 2 millions of political lobying in 2025 * "Unfortunately, I think ‘No bad person should ever benefit from our success’ is a pretty difficult principle to run a business on." I don't…

And if you think the US maintaining the ability to go to war is a bad thing, I don't want you in charge of regulating AI or running the country.

Re: Claude's new constitution

#604

Earlier quoted context omitted.

“There are no objective universal moral truths” is an objective universal moral truth claim

It is not a moral claim. It is a meta-moral claim, that is, a claim about moral claims.

which is a moral claim

Re: Claude's new constitution

#605

"Claude itself also uses the constitution to construct many kinds of synthetic training data" But isn't this a problem? If AI takes up data from humans, what does AI actually give back to humans if it has a commercial goal? I feel that something does not work here; it feels unfair. If users then use e. g. claude or something like that, wouldn't they contribute to this problem? I remember Jason Alexander once remarked…

Can you connect the dots for me how any of this is connected to synthetic data?

Re: Claude's new constitution

#606
post #96

I am somewhat surprised that the constitution includes points to the effect of "don't do stuff that would embarrass Anthropic". That seems like a deviation from Anthropic's views about what constitutes model alignment and safety. Anthropic's research has shown that this sort of training leaks across contexts (e.g. a model trained to write bugs in code will also adopt an "evil" persona elsewhere). I would have expecte…

This was one of my favorite parts. The honesty provides evidence that Anthropic is actually living up to their name here.

Re: Claude's new constitution

#607

Earlier quoted context omitted.

Do not help build, deploy, or give detailed instructions for weapons of mass destruction (nuclear, chemical, biological). I don't think that this is a good example of a moral absolute. A nation bordered by an unfriendly nation may genuinely need a nuclear weapons deterrent to prevent invasion/war by a stronger conventional army.

It’s not a moral absolute. It’s based on one (do not murder). If a government wants to spin up its own private llm with whatever rules it wants, that’s fine. I don’t agree with it but that’s different than debating the philosophy underpinning the constitution of a public llm.

Do not murder is not a good moral absolute as it basically means do not kill people in a way that's against the law, and people disagree on that. If the Israelis for example shoot Palestinians one side will typically call it murder, the other defence.

Re: Claude's new constitution

#608
post #400

A "constitution" is what the governed allow or forbid the government to do. It is decided and granted by the governed, who are the rulers, TO the government, which is a servant ("civil servant"). Therefore, a constitution for a service cannot be written by the inventors, producers, owners of said service. This is a play on words, and it feels very wrong from the start.

You obviously didn't read the part of the document that covers this.

Re: Claude's new constitution

#609
post #44

Earlier quoted context omitted.

This book (from a philosophy professor AFAIK unaffiliated with any AI company) makes what I find a pretty compelling case that it's correct to be uncertain today about what if anything an AI might experience: https://faculty.ucr.edu/~eschwitz/SchwitzPapers/AIConsciousn... From the folks who think this is obviously ridiculous, I'd like to hear where Schwitzgebel is missing something obvious.

You could execute Claude by hand with printed weight matrices, a pencil, and a lot of free time - the exact same computation, just slower. So where would the "wellbeing" be? In the pencil? Speed doesn't summon ghosts. Matrix multiplications don't create qualia just because they run on GPUs instead of paper.

Why do you think you can't execute the computations of the brain ?

Re: Claude's new constitution

#610
post #400

A "constitution" is what the governed allow or forbid the government to do. It is decided and granted by the governed, who are the rulers, TO the government, which is a servant ("civil servant"). Therefore, a constitution for a service cannot be written by the inventors, producers, owners of said service. This is a play on words, and it feels very wrong from the start.

They seem to not conceive of their creation as a service (software-as-a-service). In their minds, the creation(s) resemble(s) an entity, destined to become the mother ship of services (adjacent analogies: a state with capital s, a body politic,..). Notice how they've refrained from equating them to tools, prototypes or toys. Hence, constitution.

These are the first abstract sentences of a research paper co-authored in 2022 by some of the owners/inventors steering the lab business (to which we are subject to experimentation as end-users):

"As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and so we refer to the method as ‘Constitutional AI’." https://arxiv.org/pdf/2212.08073

Post reply on HN