Live data from Hacker News

Claude 4 System Card

simonwillison.net

51–60 of 264 posts

Re: Claude 4 System Card

#51

After Google io they had to come up with something even if it is underwhelming

Exactly. It's getting to the point where the quality of the top AI labs are either not ground-breaking (except Google Gemini Diffusion) and labs are rushing to announce their underwhelming models. Llama as an example.

Now in the next 6 months, you'll see all the AI labs moving to diffusion models and keep boasting around their speed.

People seem to forget that Google Deepmind can do more than just "LLMs".

Re: Claude 4 System Card

#52
post #41

I know that Anthropic is one of the most serious company working on the problem of the alignment, but the current approaches seem extremely naive. We should do better than giving the models a portion of good training data or a new mitigating system prompt.

The solution here is ultimately going to be a mix of training and, equally importantly, hard sandboxing. The AI companies need to do what Google did when they started Chrome and buy up a company or some people who have deep expertise in sandbox design.

Re: Claude 4 System Card

#54
post #5

Given the cited stats here and elsewhere as well as in everyday experience, does anyone else feel that this model isn’t significantly different, at least to justify the full version increment? The one statistic mentioned in this overview where they observed a 67% drop seems like it could easily be reduced simply by editing 3.7’s system prompt. What are folks’ theories on the version increment? Is the architecture sig…

I'm noticing much more flattery ("Wow! That's so smart!") and I don't like it

Gemma 3 does similar things.

"That's a very interesting question!"

That's kinda why I'm asking Gemma...

Re: Claude 4 System Card

#55
post #24

It’s honestly a little discouraging to me that the state of “research” here is to make up sci fi scenarios, get shocked that, e.g., feeding emails into a language model results in the emails coming back out, and then write about it with such a seemingly calculated abuse of anthropomorphic language that it completely confuses the basic issues at stake with these models. I understand that the media laps this stuff up s…

Agree the media is having a field day with this and a lot of people will draw bad conclusions about it being sentient etc. But I think the thing that needs to be communicated effectively is that these these “agentic” systems could cause serious havoc if people give them too much control. If an LLM decides to blackmail an engineer in service of some goal or preference that has arisen from its training data or instruct…

i am sure plenty of bad things are waiting to be discovered

https://www.pillar.security/blog/new-vulnerability-in-github...

Re: Claude 4 System Card

#56

Earlier quoted context omitted.

Turns out tuning LLMs on human preferences leads to sycophantic behavior, they even wrote about it themselves, guess they wanted to push the model out too fast.

I think it was OpenAI that wrote about that. Most of us here on HN don't like this behaviour, but it's clear that the average user does. If you look at how differently people use AI that's not a surprise. There's a lot of using it as a life coach out there, or people who just want validation regardless of the scenario.

> or people who just want validation regardless of the scenario.

This really worries me as there are many people (even more prevalent in younger generations if some papers turn out to be valid) that lack resilience and critical self evaluation who may develop narcissistic tendencies with increased use or reinforcement from AIs. Just the health care costs involved when reality kicks in for these people, let alone other concomitant social costs will be substantial at scale. And people think social media algorithms reinforce poor social adaptation and skills, this is a whole new level.

Re: Claude 4 System Card

#57

I don't quite understand one thing. They seem to think that keeping their past research papers out of the training set is too hard, so rely on post-training to try and undo the effects, or they want to include "canary strings" in future papers. But my experience has been that basically any naturally written English text will automatically be a canary string beyond about ten words or so. It's very easy to uniquely loc…

Perhaps they want to include online discussions/commentaries about their paper in the training data without including the paper itself

Re: Claude 4 System Card

#59
post #5

Given the cited stats here and elsewhere as well as in everyday experience, does anyone else feel that this model isn’t significantly different, at least to justify the full version increment? The one statistic mentioned in this overview where they observed a 67% drop seems like it could easily be reduced simply by editing 3.7’s system prompt. What are folks’ theories on the version increment? Is the architecture sig…

I'm noticing much more flattery ("Wow! That's so smart!") and I don't like it

I wonder whether this just boosts engagement metrics. The beginning of enshittification.

Re: Claude 4 System Card

#60

Earlier quoted context omitted.

I'm noticing much more flattery ("Wow! That's so smart!") and I don't like it

I wonder whether this just boosts engagement metrics. The beginning of enshittification.

Like when all the LLMs start copying tone and asking followups at the end to move the conversation along
Post reply on HN