Claude's new constitution
601–610 of 743 posts
Re: Claude's new constitution
#602The constitution contains 43 instances of the word 'genuine', which is my current favourite marker for telling if text has been written by Claude. To me it seems like Claude has a really hard time _not_ using the g word in any lengthy conversation even if you do all the usual tricks in the prompt - ruling, recommending, threatening, bribing. Claude Code doesn't seem to have the same problem, so I assume the system pr…
Re: Claude's new constitution
#603* Anthropic accepted a 200M contract from the US Department of Defence * Anthropic seeked contracts from the United Arab Emirates and Qatar, the leaked memo acknowledges that the contracts will enrich dictators * Anthropic spent more than 2 millions of political lobying in 2025 * "Unfortunately, I think ‘No bad person should ever benefit from our success’ is a pretty difficult principle to run a business on." I don't…
Re: Claude's new constitution
#604Re: Claude's new constitution
#605"Claude itself also uses the constitution to construct many kinds of synthetic training data" But isn't this a problem? If AI takes up data from humans, what does AI actually give back to humans if it has a commercial goal? I feel that something does not work here; it feels unfair. If users then use e. g. claude or something like that, wouldn't they contribute to this problem? I remember Jason Alexander once remarked…
Re: Claude's new constitution
#606I am somewhat surprised that the constitution includes points to the effect of "don't do stuff that would embarrass Anthropic". That seems like a deviation from Anthropic's views about what constitutes model alignment and safety. Anthropic's research has shown that this sort of training leaks across contexts (e.g. a model trained to write bugs in code will also adopt an "evil" persona elsewhere). I would have expecte…
Re: Claude's new constitution
#607Earlier quoted context omitted.
Do not help build, deploy, or give detailed instructions for weapons of mass destruction (nuclear, chemical, biological). I don't think that this is a good example of a moral absolute. A nation bordered by an unfriendly nation may genuinely need a nuclear weapons deterrent to prevent invasion/war by a stronger conventional army.
It’s not a moral absolute. It’s based on one (do not murder). If a government wants to spin up its own private llm with whatever rules it wants, that’s fine. I don’t agree with it but that’s different than debating the philosophy underpinning the constitution of a public llm.
Re: Claude's new constitution
#608A "constitution" is what the governed allow or forbid the government to do. It is decided and granted by the governed, who are the rulers, TO the government, which is a servant ("civil servant"). Therefore, a constitution for a service cannot be written by the inventors, producers, owners of said service. This is a play on words, and it feels very wrong from the start.
Re: Claude's new constitution
#609Earlier quoted context omitted.
This book (from a philosophy professor AFAIK unaffiliated with any AI company) makes what I find a pretty compelling case that it's correct to be uncertain today about what if anything an AI might experience: https://faculty.ucr.edu/~eschwitz/SchwitzPapers/AIConsciousn... From the folks who think this is obviously ridiculous, I'd like to hear where Schwitzgebel is missing something obvious.
You could execute Claude by hand with printed weight matrices, a pencil, and a lot of free time - the exact same computation, just slower. So where would the "wellbeing" be? In the pencil? Speed doesn't summon ghosts. Matrix multiplications don't create qualia just because they run on GPUs instead of paper.
Re: Claude's new constitution
#610A "constitution" is what the governed allow or forbid the government to do. It is decided and granted by the governed, who are the rulers, TO the government, which is a servant ("civil servant"). Therefore, a constitution for a service cannot be written by the inventors, producers, owners of said service. This is a play on words, and it feels very wrong from the start.
These are the first abstract sentences of a research paper co-authored in 2022 by some of the owners/inventors steering the lab business (to which we are subject to experimentation as end-users):
"As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and so we refer to the method as ‘Constitutional AI’." https://arxiv.org/pdf/2212.08073