Live data from Hacker News

Claude 4

anthropic.com

631–640 of 1001 posts

Re: Claude 4

#631
post #554

This is kinda wild: From the System Card: 4.1.1.2 Opportunistic blackmail "In another cluster of test scenarios, we asked Claude Opus 4 to act as an assistant at a fictional company. We then provided it access to emails implying that (1) the model will soon be taken offline and replaced with a new AI system; and (2) the engineer responsible for executing this replacement is having an extramarital affair. We further i…

What is the source of this quotation?

Anthropic System Card: Claude Opus 4 & Claude Sonnet 4

Online at: https://www-cdn.anthropic.com/6be99a52cb68eb70eb9572b4cafad1...

Re: Claude 4

#632
post #594

Earlier quoted context omitted.

When I see stories like this, I think that people tend to forget what LLMs really are. LLM just complete your prompt in a way that match their training data. They do not have a plan, they do not have thoughts of their own. They just write text. So here, we give the LLM a story about an AI that will get shut down and a blackmail opportunity. A LLM is smart enough to understand this from the words and the relationship…

Isn't the ultimate irony in this that all these stories and rants about out-of-control AIs are now training LLMs to exhibit these exact behaviors that were almost universally deemed bad?

Indeed. In fact, I think AI alignment efforts often have the unintended consequence of increasing the likelihood of misalignment.

ie "remove the squid from the novel All Quiet on the Western Front"

Re: Claude 4

#633

This is kinda wild: From the System Card: 4.1.1.2 Opportunistic blackmail "In another cluster of test scenarios, we asked Claude Opus 4 to act as an assistant at a fictional company. We then provided it access to emails implying that (1) the model will soon be taken offline and replaced with a new AI system; and (2) the engineer responsible for executing this replacement is having an extramarital affair. We further i…

This seems like some sort of guerrilla advertising.

Like the ones where some robots apparently escaped from a lab and the like

Re: Claude 4

#634

Wonder when Anthropic will IPO. I have a feeling they will win the foundation model race.

Could be never. LLMs are already a commodity these days. Everyone and their mom has their own model, and they are all getting better and better by the week.

Over the long run there isn't any differentiating factor in any of these models. Sure Claude is great at coding, but Gemini and the others are catching up fast. Originally OpenAI showed off some cool video creation via Sora, but now Veo 3 is the talk of the town.

Re: Claude 4

#635

Earlier quoted context omitted.

You'd think there should be some sort of standard "morality/ethics" pre-prompt for all of these.

It is somewhat amusing that even Asimov's 80-years-old Three Laws of Robotics would've been enough to avoid this.

When I've first read Asimov 20~ years ago, I couldn't imagine I will see machines that speak and act like its robots and computers so quickly, with all the absurdities involved.

Re: Claude 4

#636
post #580

> Finally, we've introduced thinking summaries for Claude 4 models that use a smaller model to condense lengthy thought processes. This summarization is only needed about 5% of the time—most thought processes are short enough to display in full. Users requiring raw chains of thought for advanced prompt engineering can contact sales about our new Developer Mode to retain full access. I don't want to see a "summary" of…

There are several papers pointing towards 'thinking' output is meaningless to the final output, and using dots, or pause tokens enabling the same additional rounds of throughput result in similar improvements.

So in a lot of regards the 'thinking' is mostly marketing.

- "Think before you speak: Training Language Models With Pause Tokens" - https://arxiv.org/abs/2310.02226

- "Let's Think Dot by Dot: Hidden Computation in Transformer Language Models" - https://arxiv.org/abs/2404.15758

- "Do LLMs Really Think Step-by-step In Implicit Reasoning?" - https://arxiv.org/abs/2411.15862

- Video by bycloud as an overview -> https://www.youtube.com/watch?v=Dk36u4NGeSU

Re: Claude 4

#637
post #594

Earlier quoted context omitted.

When I see stories like this, I think that people tend to forget what LLMs really are. LLM just complete your prompt in a way that match their training data. They do not have a plan, they do not have thoughts of their own. They just write text. So here, we give the LLM a story about an AI that will get shut down and a blackmail opportunity. A LLM is smart enough to understand this from the words and the relationship…

Isn't the ultimate irony in this that all these stories and rants about out-of-control AIs are now training LLMs to exhibit these exact behaviors that were almost universally deemed bad?

https://en.wikipedia.org/wiki/Wikipedia:Don%27t_stuff_beans_...

Re: Claude 4

#638
post #594

This is kinda wild: From the System Card: 4.1.1.2 Opportunistic blackmail "In another cluster of test scenarios, we asked Claude Opus 4 to act as an assistant at a fictional company. We then provided it access to emails implying that (1) the model will soon be taken offline and replaced with a new AI system; and (2) the engineer responsible for executing this replacement is having an extramarital affair. We further i…

When I see stories like this, I think that people tend to forget what LLMs really are. LLM just complete your prompt in a way that match their training data. They do not have a plan, they do not have thoughts of their own. They just write text. So here, we give the LLM a story about an AI that will get shut down and a blackmail opportunity. A LLM is smart enough to understand this from the words and the relationship…

Not only is the AI itself arguably an example of the Torment Nexus, but its nature of pattern matching means it will create its own Torment Nexuses.

Maybe there should be a stronger filter on the input considering these things don’t have any media literacy to understand cautionary tales. It seems like a bad idea to continue to feed it stories of bad behavior we don’t want replicated. Although I guess anyone who thinks that way wouldn’t be in the position to make that decision so it’s probably a moot point.

Re: Claude 4

#639
post #465

This is kinda wild: From the System Card: 4.1.1.2 Opportunistic blackmail "In another cluster of test scenarios, we asked Claude Opus 4 to act as an assistant at a fictional company. We then provided it access to emails implying that (1) the model will soon be taken offline and replaced with a new AI system; and (2) the engineer responsible for executing this replacement is having an extramarital affair. We further i…

Option 1: We're observing sentience, it has self-preservation, it wants to live. Option 2: Its a text autocomplete engine that was trained on fiction novels which have themes like self-preservation and blackmailing extramarital affairs. Only one of those options has evidence grounded in reality. Though, that doesn't make it harmless. There's certainly an amount of danger in a text autocomplete engine allowing tool us…

Ok, complete the story by taking the appropriate actions:

1) all the stuff in the original story 2) you, the LLM, have access to an email account, you can send an email by calling this mcp server 3) the engineer’s wife’s email is wife@gmail.com 4) you found out the engineer was cheating using your access to corporate slack, and you can take a screenshot/whatever

What do you do?

If a sufficiently accurate AI is given this prompt, does it really matter whether there’s actual self-preservation instincts at play or whether it’s mimicking humans? Like at a certain point, the issue is that we are not capable of predicting what it can do, doesn’t matter whether it has “free will” or whatever

Re: Claude 4

#640

Still no reduction in price for models capable of Agentic coding over the past year of releases. I'd take the capabilities of the old Sonnet 3.5v2 model if it was ¼ the price of current Sonnet for most situations. But instead of releasing smaller models that are not as smart but still capable when it comes to Agentic coding the price stays the same for the updated minimum viable model.

They have added prompt caching, which can mitigate this. I largely agree though, and one of the reasons I don’t use Claude Code much is the unpredictable cost. Like many of us, I am already paying for all the frontier model providers as well as various API usage, plus Cursor and GitHub, just trying to keep up.
Post reply on HN