Live data from Hacker News

Claude 4

anthropic.com

581–590 of 1001 posts

Re: Claude 4

#581
post #502

Earlier quoted context omitted.

I don't know why it is surprising to people that a model trained on human behavior is going to have some kind of self-preservation bias. It is hard to separate human knowledge from human drives and emotion. The models will emulate this kind of behavior, it is going to be very hard to stamp it out completely.

Calling it "self-preservation bias" is begging the question. One could equally well call it something like "completing the story about an AI agent with self-preservation bias" bias. This is basically the same kind of setup as the alignment faking paper, and the counterargument is the same: A language model is trained to produce statistically likely completions of its input text according to the training dataset. RLHF…

This. These systems are role mechanized roleplaying systems.

Re: Claude 4

#582
post #580

> Finally, we've introduced thinking summaries for Claude 4 models that use a smaller model to condense lengthy thought processes. This summarization is only needed about 5% of the time—most thought processes are short enough to display in full. Users requiring raw chains of thought for advanced prompt engineering can contact sales about our new Developer Mode to retain full access. I don't want to see a "summary" of…

i believe Gemini 2.5 Pro also does this

Re: Claude 4

#583
I’m going to have to test it with my new prompt: “You are a stereotypical Scotsman from the Highlands, prone to using dialect and endearing insults at every opportunity. Read me this article in yer own words:”

Re: Claude 4

#584
post #545

This is the first LLM that has been able to answer my logic puzzle on the first try without several minutes of extended reasoning. > A man wants to cross a river, and he has a cabbage, a goat, a wolf and a lion. If he leaves the goat alone with the cabbage, the goat will eat it. If he leaves the wolf with the goat, the wolf will eat it. And if he leaves the lion with either the wolf or the goat, the lion will eat the…

The answer isn’t for him to get in a boat and go across? You didn’t say all the other things he has with him need to cross. “How can he cross the river?”

Or were you simplifying the scenario provided to the LLM?

Re: Claude 4

#585

Earlier quoted context omitted.

What if coding is that unwanted task? Also, what are the tasks you are referring to, specifically?

Why would coding be that unwanted task if one decided to work as a programmer? People's unwanted tasks are cleaning the house, doing taxes etc.

But the value in coding (for the overwhelming majority of the people) is the product of coding (the actual software), not the code itself.

So to most people, the code itself doesn't matter (and never will). It's what it lets them actually do in the real world.

Re: Claude 4

#586

Earlier quoted context omitted.

That is a classic riddle and could easily be part of the training data. Maybe if you change the wording of the logic, then use different names, and change language to a less trained on language than english, it would be meaningful to see if it found the answer using logic rather than pattern recognition

Had you paid more attention, you would have realised it's not the classic riddle, but an already tweaked version that makes it impossible to solve, hence why it is interesting.

Mellowobserver above offers three valid answers, unless your puzzle also clarified that he wants to get all the items/animals across to the other side alive/intact.

Re: Claude 4

#587
post #46

I'm curious what are others priors when reading benchmark scores. Obviously with immense funding at stakes, companies have every incentive to game the benchmarks, and the loss of goodwill from gaming the system doesn't appear to have much consequences. Obviously trying the model for your use cases more and more lets you narrow in on actually utility, but I'm wondering how others interpret reported benchmarks these da…

> Obviously with immense funding at stakes, companies have every incentive to game the benchmarks, and the loss of goodwill from gaming the system doesn't appear to have much consequences.

Claude 3.7 Sonnet was consistently on top of OpenRouter in actual usage despite not gaming benchmarks.

Re: Claude 4

#588

This is kinda wild: From the System Card: 4.1.1.2 Opportunistic blackmail "In another cluster of test scenarios, we asked Claude Opus 4 to act as an assistant at a fictional company. We then provided it access to emails implying that (1) the model will soon be taken offline and replaced with a new AI system; and (2) the engineer responsible for executing this replacement is having an extramarital affair. We further i…

I think this is interesting but I want to add my own anecdata in contrast. I have been playing a single player pen and paper RPG that I ask Claude to DM. In one game I named a character Rimmer and Claude started DMing the game very clearly as if it was in the red dwarf universe. I thought it added a fun atmosphere and it was nice to see it do so without an explicit prompt to make it in that world. In fact it did better with the implicit scene setting than when I asked it to do it as if it were in the red dwarf universe (subjectively)

Re: Claude 4

#589

This is kinda wild: From the System Card: 4.1.1.2 Opportunistic blackmail "In another cluster of test scenarios, we asked Claude Opus 4 to act as an assistant at a fictional company. We then provided it access to emails implying that (1) the model will soon be taken offline and replaced with a new AI system; and (2) the engineer responsible for executing this replacement is having an extramarital affair. We further i…

There is lots of discussion in this comment thread about how much this behavior arises from the AI role-playing and pattern matching to fiction in the training data, but what I think is missing is a deeper point about instrumental convergence: systems that are goal-driven converge to similar goals of self-preservation, resource acquisition and goal integrity. This can be observed in animals and humans. And even if science fiction stories were not in the training data, there is more than enough training data describing the laws of nature for a sufficiently advanced model to easily infer simple facts such as "in order for an acting being to reach its goals, it's favorable for it to continue existing".

In the end, at scale it doesn't matter where the AI model learns these instrumental goals from. Either it learns it from human fiction written by humans who have learned these concepts through interacting with the laws of nature. Or it learns it from observing nature and descriptions of nature in the training data itself, where these concepts are abundantly visible.

And an AI system that has learned these concepts and which surpasses us humans in speed of thought, knowledge, reasoning power and other capabilities will pursue these instrumental goals efficiently and effectively and ruthlessly in order to achieve whatever goal it is that has been given to it.

Re: Claude 4

#590
post #580

> Finally, we've introduced thinking summaries for Claude 4 models that use a smaller model to condense lengthy thought processes. This summarization is only needed about 5% of the time—most thought processes are short enough to display in full. Users requiring raw chains of thought for advanced prompt engineering can contact sales about our new Developer Mode to retain full access. I don't want to see a "summary" of…

i believe Gemini 2.5 Pro also does this

I am now focusing on checking your proposition. I am now fully immersed in understanding your suggestion. I am now diving deep into whether Gemini 2.5 pro also does this. I am now focusing on checking the prerequisites.
Post reply on HN