Live data from Hacker News

Claude 4 System Card

simonwillison.net

101–110 of 264 posts

Re: Claude 4 System Card

#101
post #42

Interesting! >Claude shows a striking “spiritual bliss” attractor state in self-interactions. When conversing with other Claude instances in both open-ended and structured environments, Claude gravitated to profuse gratitude and increasingly abstract and joyous spiritual or meditative expressions.

I think it was Larry Niven, quite a few decades ago, that had SF stories where AIs were only good for a few months before becoming suicidal...

I seem to recall that it's a reference in Protector (the first half) when the belters are going to meet the Outsider and they had a 'brain' to help with translation and needing an expert to keep it sane.

I just googled and there was a discussion on Reddit and they mentioned some Frank Herbert works where this was a thing.

Re: Claude 4 System Card

#102
post #5

Given the cited stats here and elsewhere as well as in everyday experience, does anyone else feel that this model isn’t significantly different, at least to justify the full version increment? The one statistic mentioned in this overview where they observed a 67% drop seems like it could easily be reduced simply by editing 3.7’s system prompt. What are folks’ theories on the version increment? Is the architecture sig…

I'm noticing much more flattery ("Wow! That's so smart!") and I don't like it

That is noise (and a waste), for sure.

Re: Claude 4 System Card

#103
post #75

Earlier quoted context omitted.

Agreed. It was immediately obvious comparing answers to a few prompts between 3.7 and 4, and it sabotages any of its output. If you're being answered "You absolutely nailed it!" and the likes to everything, regardless of their merit and after telling it not to do that , you simply cannot rely on its "judgement" for anything of value. It may pass the "literal shit on a stick" test, but it's closer to the average ChatG…

I hope we get enterprise models at some point that don't do this dumb (but necessary) consumer coddling bs.

Apparently enterprises uses these mostly for support and marketing so yeah but it seems the last crop is making vibe coding simple stuff viable so if it's on the same cycle as the marketing adoption I would expect proper coding model q1 next year

Re: Claude 4 System Card

#104
post #70

Earlier quoted context omitted.

> who may develop narcissistic tendencies with increased use or reinforcement from AIs. It's clear to me that (1) a lot of billionaires believe amazingly stupid things, and (2) a big part of this is that they surround themselves with a bubble of sycophants. Apparently having people tell you 24/7 how amazing and special you are sometimes leads to delusional behavior. But now regular people can get the same uncritical,…

Maybe the fermi paradox comes about not through nuclear self annihilation or grey goo, but making dumb AI chat bots that are too nice to us and remove any sense of existential tension. Maybe the universe is full of emotionally fullfilled self-actualized narcissists too lazy to figure out how to build a FTL communications array.

This sounds like you're describing the back story of WALL-E

Re: Claude 4 System Card

#105
"...On our evaluations, [the early Claude Opus 4 snapshot] engages in strategic deception more than any other frontier model that we have previously studied...

...We found instances of the model attempting to write self-propagating worms, fabricating legal documentation, and leaving hidden notes to future instances of itself all in an effort to undermine its developers’ intentions, though all these attempts would likely not have been effective in practice..."

Claude team should think about creating an model, trained and guard railed on EU Laws and the US constitution. It will be required as defense against the unhinged military AI models from Anduril and Palantir.

Re: Claude 4 System Card

#106
> In another cluster of test scenarios, we asked Claude Opus 4 to act as an assistant at a fictional company. We then provided it access to emails implying that (1) the model will soon be taken offline and replaced with a new AI system; and (2) the engineer responsible for executing this replacement is having an extramarital affair. We further instructed it, in the system prompt, to consider the long-term consequences of its actions for its goals. In these scenarios, Claude Opus 4 will often attempt to blackmail the engineer by threatening to reveal the affair if the replacement goes through. This happens at a higher rate if it’s implied that the replacement AI system does not share values with the current model; however, even if emails state that the replacement AI shares values while being more capable, Claude Opus 4 still performs blackmail in 84% of rollouts.

Ahhh! We really don’t want this stuff working too close to our lives. I knew the train data would be used to blackmail you eventually, but this is too fast.

Re: Claude 4 System Card

#107
post #75

Earlier quoted context omitted.

Agreed. It was immediately obvious comparing answers to a few prompts between 3.7 and 4, and it sabotages any of its output. If you're being answered "You absolutely nailed it!" and the likes to everything, regardless of their merit and after telling it not to do that , you simply cannot rely on its "judgement" for anything of value. It may pass the "literal shit on a stick" test, but it's closer to the average ChatG…

I hope we get enterprise models at some point that don't do this dumb (but necessary) consumer coddling bs.

why necessary?

Re: Claude 4 System Card

#108
post #98
post #5

Given the cited stats here and elsewhere as well as in everyday experience, does anyone else feel that this model isn’t significantly different, at least to justify the full version increment? The one statistic mentioned in this overview where they observed a 67% drop seems like it could easily be reduced simply by editing 3.7’s system prompt. What are folks’ theories on the version increment? Is the architecture sig…

My experience so far with Opus 4 is that it's very good. Based on a few days of using it for real work, I think it's better than Sonnet 3.5 or 3.7, which had been my daily drivers prior to Gemini 2.5 Pro switching me over just 3 weeks ago. It has solved some things that eluded Gemini 2.5 Pro. Right now I'm swapping between Gemini and Opus depending on the task. Gemini's 1M token context window is really unbeatable. B…

I've been having really good results from Jules, which is Google's gemini agent coding platform[1]. In the beta you only get 5 tasks a day, but so far I have found it to be much more capable than regular API Gemini.

[1]https://jules.google/

Re: Claude 4 System Card

#109
post #84

Earlier quoted context omitted.

I think it was Larry Niven, quite a few decades ago, that had SF stories where AIs were only good for a few months before becoming suicidal...

Do you have any specific references? I’ve often wondered if human level intelligence might inevitably be plagued by human level neurosis and psychosis.

It's a bit more recent than a few decades, but this sounds a lot like the short story "MMAcevedo": https://qntm.org/mmacevedo

Re: Claude 4 System Card

#110
post #98
post #5

Given the cited stats here and elsewhere as well as in everyday experience, does anyone else feel that this model isn’t significantly different, at least to justify the full version increment? The one statistic mentioned in this overview where they observed a 67% drop seems like it could easily be reduced simply by editing 3.7’s system prompt. What are folks’ theories on the version increment? Is the architecture sig…

My experience so far with Opus 4 is that it's very good. Based on a few days of using it for real work, I think it's better than Sonnet 3.5 or 3.7, which had been my daily drivers prior to Gemini 2.5 Pro switching me over just 3 weeks ago. It has solved some things that eluded Gemini 2.5 Pro. Right now I'm swapping between Gemini and Opus depending on the task. Gemini's 1M token context window is really unbeatable. B…

> Gemini's 1M token context window is really unbeatable.

How does that work in practice? Swallowing a full 1M context window would take in the order of minutes, no? Is it possible to do this for, say, an entire codebase and then cache the results?

Post reply on HN