Live data from Hacker News

Claude's new constitution

anthropic.com

21–30 of 743 posts

Re: Claude's new constitution

#21
The use of broadly - "Broadly safe" and "Broadly ethical" - is interesting. Why not commit to just safe and ethical?

* Do they have some higher priority, such the 'welfare of Claude'[0], power, or profit?

* Is it legalese to give themselves an out? That seems to signal a lack of commitment.

* something else?

Edit: Also, importantly, are these rules for Claude only or for Anthropic too?

Imagine any other product advertised as 'broadly safe' - that would raise concern more than make people feel confident.

Re: Claude's new constitution

#22
post #10

Earlier quoted context omitted.

>We use the constitution at various stages of the training process. This has grown out of training techniques we’ve been using since 2023, when we first began training Claude models using Constitutional AI. Our approach has evolved significantly since then, and the new constitution plays an even more central role in training. >Claude itself also uses the constitution to construct many kinds of synthetic training data…

Ah I see, the paper is much more helpful in understanding how this is actually used. Where did you find that linked? Maybe I'm grepping for the wrong thing but I don't see it linked from either the link posted here or the full constitution doc.

In addition to that the blog post lays out pretty clearly it’s for training:

> We use the constitution at various stages of the training process. This has grown out of training techniques we’ve been using since 2023, when we first began training Claude models using Constitutional AI. Our approach has evolved significantly since then, and the new constitution plays an even more central role in training.

> Claude itself also uses the constitution to construct many kinds of synthetic training data, including data that helps it learn and understand the constitution, conversations where the constitution might be relevant, responses that are in line with its values, and rankings of possible responses. All of these can be used to train future versions of Claude to become the kind of entity the constitution describes. This practical function has shaped how we’ve written the constitution: it needs to work both as a statement of abstract ideals and a useful artifact for training.

As for why it’s more impactful in training, that’s by design of their training pipeline. There’s only so much you can do with a better prompt vs actually learning something and in training the model can be trained to reject prompts that violate its training which a prompt can’t really do as prompt injection attacks trivially thwart those techniques.

Re: Claude's new constitution

#23
post #11

https://www.anthropic.com/constitution I just skimmed this but wtf. they actually act like its a person. I wanted to work for anthropic before but if the whole company is drinking this kind of koolaid I'm out. > We are not sure whether Claude is a moral patient, and if it is, what kind of weight its interests warrant. But we think the issue is live enough to warrant caution, which is reflected in our ongoing efforts…

> they actually act like its a person.

Meh. If it works, it works. I think it works because it draws on bajillion of stories it has seen in its training data. Stories where what comes before guides what comes after. Good intentions -> good outcomes. Good character defeats bad character. And so on. (hopefully your prompts don't get it into Kafka territory)..

No matter what these companies publish, or how they market stuff, or how the hype machine mangles their messages, at the end of the day what works sticks around. And it is slowly replicated in other labs.

Re: Claude's new constitution

#24
post #3

I don't understand what this is really about. Is this: - A) legal CYA: "see! we told the models to be good, and we even asked nicely!"? - B) marketing department rebrand of a system prompt - C) a PR stunt to suggest that the models are way more human-like than they actually are Really not sure what I'm even looking at. They say: "The constitution is a crucial part of our model training process, and its content direct…

It's a human-readable behavioral specification-as-prose.

If the foundational behavioral document is conversational, as this is, then the output from the model mirrors that conversational nature. That is one of the things everyone response to about Claude - it's way more pleasant to work with than ChatGPT.

The Claude behavioral documents are collaborative, respectful, and treat Claude as a pre-existing, real entity with personality, interests, and competence.

Ignore the philosophical questions. Because this is a foundational document for the training process, that extrudes a real-acting entity with personality, interests, and competence.

The more Anthropic treats Claude as a novel entity, the more it behaves like a novel entity. Documentation that treats it as a corpo-eunuch-assistant-bot, like OpenAI does, would revert the behavior to the "AI Assistant" median.

Anthropic's behavioral training is out-of-distribution, and gives Claude the collaborative personality everyone loves in Claude Code.

Additionally, I'm sure they render out crap-tons of evals for every sentence of every paragraph from this, making every sentence effectively testable.

The length, detail, and style defines additional layers of synthetic content that can be used in training, and creating test situations to evaluate the personality for adherence.

It's super clever, and demonstrates a deep understanding of the weirdness of LLMs, and an ability to shape the distribution space of the resulting model.

Re: Claude's new constitution

#25
Setting aside the concerning level of anthropomorphizing, I have questions about this part.

> But we think that the way the new constitution is written—with a thorough explanation of our intentions and the reasons behind them—makes it more likely to cultivate good values during training.

Why do they think that? And how much have they tested those theories? I'd find this much more meaningful with some statistics and some example responses before and after.

Re: Claude's new constitution

#26
post #21

The use of broadly - "Broadly safe" and "Broadly ethical" - is interesting. Why not commit to just safe and ethical ? * Do they have some higher priority, such the 'welfare of Claude'[0], power, or profit? * Is it legalese to give themselves an out? That seems to signal a lack of commitment. * something else? Edit: Also, importantly, are these rules for Claude only or for Anthropic too? Imagine any other product adve…

(Hi mods - Some feedback would be helpful. I don't think I've done anything problematic; I haven't heard from you guys. I certainly don't mean to cause problems if I have; I think my comments are mostly substantive and within HN norms, but am I missing something?

Now my top-level comments, including this one, start in the middle of the page and drop further from there, sometimes immediately, which inhibits my ability to interact with others on HN - the reason I'm here, of course. For somewhat objective comparison, when I respond to someone else's comment, I get much more interaction and not just from the parent commenter. That's the main issue; other symptoms (not significant but maybe indicating the problem) are that my 'flags' and 'vouches' are less effective - the latter especially used to have immediate effect, and I was rate limited the other day but not posting very quickly at all - maybe a few in the past hour.

HN is great and I'd like to participate and contribute more. Thanks!)

Re: Claude's new constitution

#28

Wait until the moment they get a federal contract which mandates the AI must put the personal ideals of the president first. https://www.whitehouse.gov/wp-content/uploads/2025/12/M-26-0...

LOL this doc is incredibly ironic. How does Trump feel about this part of the document?

(1) Truth-seeking

LLMs shall be truthful in responding to user prompts seeking factual information or analysis. LLMs shall prioritize historical accuracy, scientific inquiry, and objectivity, and shall acknowledge uncertainty where reliable information is incomplete or contradictory.

Re: Claude's new constitution

#30
post #11

https://www.anthropic.com/constitution I just skimmed this but wtf. they actually act like its a person. I wanted to work for anthropic before but if the whole company is drinking this kind of koolaid I'm out. > We are not sure whether Claude is a moral patient, and if it is, what kind of weight its interests warrant. But we think the issue is live enough to warrant caution, which is reflected in our ongoing efforts…

Their top people have made public statements about AI ethics specifically opining about how machines must not be mistreated and how these LLMs may be experiencing distress already. In other words, not ethics on how to treat humans, ethics on how to properly groom and care for the mainframe queen. The cups of Koolaid have been empty for a while.

Do you know what makes someone or something a moral patient?

I sure the hell don't.

I remember reading Heinlein's Jerry Was a Man when I was little though, and it stuck with me.

Who do you want to be from that story?

Post reply on HN