Live data from Hacker News

Claude 4

anthropic.com

611–620 of 1001 posts

Re: Claude 4

#611
post #465

This is kinda wild: From the System Card: 4.1.1.2 Opportunistic blackmail "In another cluster of test scenarios, we asked Claude Opus 4 to act as an assistant at a fictional company. We then provided it access to emails implying that (1) the model will soon be taken offline and replaced with a new AI system; and (2) the engineer responsible for executing this replacement is having an extramarital affair. We further i…

Option 1: We're observing sentience, it has self-preservation, it wants to live. Option 2: Its a text autocomplete engine that was trained on fiction novels which have themes like self-preservation and blackmailing extramarital affairs. Only one of those options has evidence grounded in reality. Though, that doesn't make it harmless. There's certainly an amount of danger in a text autocomplete engine allowing tool us…

The only proof that anyone is sentient is that you experience sentience and assume others are sentient because they are similar to you.

On a practical level there is no difference between a sentient being, and a machine that is extremely good at role playing being sentient.

Re: Claude 4

#612
post #599

Tried Sonnet with 5-disk towers of Hanoi puzzle. Failed miserably :/ https://claude.ai/share/6afa54ce-a772-424e-97ed-6d52ca04de28

Sonnet with extended thinking solved it after 30s for me: https://claude.ai/share/b974bd96-91f4-4d92-9aa8-7bad964e9c5a Normal Opus solved it: https://claude.ai/share/a1845cc3-bb5f-4875-b78b-ee7440dbf764 Opus with extended thinking solved it after 7s: https://claude.ai/share/0cf567ab-9648-4c3a-abd0-3257ed4fbf59 Though it's a weird puzzle to use a benchmark because the answer is so formulaic.

It is formulaic which is why it surprised me that Sonnet failed it. I don't have access to the other models so I'll stick with Gemini for now.

Re: Claude 4

#613
post #250

Interesting alignment notes from Opus 4: https://x.com/sleepinyourhat/status/1925593359374328272 "Be careful about telling Opus to ‘be bold’ or ‘take initiative’ when you’ve given it access to real-world-facing tools...If it thinks you’re doing something egregiously immoral, for example, like faking data in a pharmaceutical trial, it will use command-line tools to contact the press, contact regulators, try to lock yo…

https://x.com/sleepinyourhat/status/1925626079043104830 "I deleted the earlier tweet on whistleblowing as it was being pulled out of context. TBC: This isn't a new Claude feature and it's not possible in normal usage. It shows up in testing environments where we give it unusually free access to tools and very unusual instructions."

Trying to imagine proudly bragging about my hallucination machine’s ability to call the cops and then having to assure everyone that my hallucination machine won’t call the cops but the first part makes me laugh so hard that I’m crying so I can’t even picture the second part

Re: Claude 4

#614

This is kinda wild: From the System Card: 4.1.1.2 Opportunistic blackmail "In another cluster of test scenarios, we asked Claude Opus 4 to act as an assistant at a fictional company. We then provided it access to emails implying that (1) the model will soon be taken offline and replaced with a new AI system; and (2) the engineer responsible for executing this replacement is having an extramarital affair. We further i…

Funny coincidence I'm just replaying Fallout 4 and just yesterday I followed a "mysterious signal" to the New England Technocrat Society, where all members had been killed and turned into Ghouls. What happened was that they got an AI to run the building and the AI was then aggressively trained on Horror movies etc. to prepare it to organise the upcoming Halloween party, and it decided that death and torture is what humans liked.

This seems awfully close to the same sort of scenario.

Re: Claude 4

#615
post #595
post #580

> Finally, we've introduced thinking summaries for Claude 4 models that use a smaller model to condense lengthy thought processes. This summarization is only needed about 5% of the time—most thought processes are short enough to display in full. Users requiring raw chains of thought for advanced prompt engineering can contact sales about our new Developer Mode to retain full access. I don't want to see a "summary" of…

Don't be so concerned. There's ample evidence that thinking is often disassociated from the output. My take is that this is a user experience improvement, given how little people actually goes on to read the thinking process.

then provide it as an option?

Re: Claude 4

#616

An important note not mentioned in this announcement is that Claude 4's training cutoff date is March 2025, which is the latest of any recent model. (Gemini 2.5 has a cutoff of January 2025) https://docs.anthropic.com/en/docs/about-claude/models/overv...

Should we not necessarily assume that it would have some FastHTML training with that March 2025 cutoff date? I'd hope so but I guess it's more likely that it still hasn't trained on FastHTML?

Claude 4 actually knows FastHTML pretty well! :D It managed to one-shot most basic tasks I sent its way, although it makes a lot of standard minor n00b mistakes that make its code a bit longer and more complex than needed.

I've nearly finished writing a short guide which, when added to a prompt, gives quite idiomatic FastHTML code.

Re: Claude 4

#617

Earlier quoted context omitted.

It's closer to <30k before performance degrades too much for 3.5/3.7. 200k/64k is meaningless in this context.

Is there a benchmark to measure real effective context length? Sure, gpt-4o has a context window of 128k, but it loses a lot from the beginning/middle.

ruler https://arxiv.org/abs/2404.06654

nolima https://arxiv.org/abs/2502.05167

Re: Claude 4

#619
post #594

This is kinda wild: From the System Card: 4.1.1.2 Opportunistic blackmail "In another cluster of test scenarios, we asked Claude Opus 4 to act as an assistant at a fictional company. We then provided it access to emails implying that (1) the model will soon be taken offline and replaced with a new AI system; and (2) the engineer responsible for executing this replacement is having an extramarital affair. We further i…

When I see stories like this, I think that people tend to forget what LLMs really are. LLM just complete your prompt in a way that match their training data. They do not have a plan, they do not have thoughts of their own. They just write text. So here, we give the LLM a story about an AI that will get shut down and a blackmail opportunity. A LLM is smart enough to understand this from the words and the relationship…

"They do not have a plan"

Not necessarily correct if you consider agent architectures where one LLM would come up with a plan and another LLM executes the provided plan. This is already existing.

Re: Claude 4

#620
post #594

This is kinda wild: From the System Card: 4.1.1.2 Opportunistic blackmail "In another cluster of test scenarios, we asked Claude Opus 4 to act as an assistant at a fictional company. We then provided it access to emails implying that (1) the model will soon be taken offline and replaced with a new AI system; and (2) the engineer responsible for executing this replacement is having an extramarital affair. We further i…

When I see stories like this, I think that people tend to forget what LLMs really are. LLM just complete your prompt in a way that match their training data. They do not have a plan, they do not have thoughts of their own. They just write text. So here, we give the LLM a story about an AI that will get shut down and a blackmail opportunity. A LLM is smart enough to understand this from the words and the relationship…

Isn't the ultimate irony in this that all these stories and rants about out-of-control AIs are now training LLMs to exhibit these exact behaviors that were almost universally deemed bad?
Post reply on HN