Live data from Hacker News

Claude 4 System Card

simonwillison.net

151–160 of 264 posts

Re: Claude 4 System Card

#151
post #24

It’s honestly a little discouraging to me that the state of “research” here is to make up sci fi scenarios, get shocked that, e.g., feeding emails into a language model results in the emails coming back out, and then write about it with such a seemingly calculated abuse of anthropomorphic language that it completely confuses the basic issues at stake with these models. I understand that the media laps this stuff up s…

It's a massive hype bubble unrivaled in scale by anything that has ever come before it, so all the AI providers have huge vested interests in making it seem like these systems are "sentient". All of the marketing is riddled with anthropomorphization (is that a word?). "It's like a Junior!", "It's like your secretary!", "But humans also do X!" etc. The other day on the Claude 4 announcement post [1], people were talki…

[deleted]

Re: Claude 4 System Card

#152
post #75

Earlier quoted context omitted.

I'm noticing much more flattery ("Wow! That's so smart!") and I don't like it

Agreed. It was immediately obvious comparing answers to a few prompts between 3.7 and 4, and it sabotages any of its output. If you're being answered "You absolutely nailed it!" and the likes to everything, regardless of their merit and after telling it not to do that , you simply cannot rely on its "judgement" for anything of value. It may pass the "literal shit on a stick" test, but it's closer to the average ChatG…

GPT 4o is unbearable in this sense, but o3 has very much toned it down in my experience. I don't need to wrap my prompts or anything.

Re: Claude 4 System Card

#153
post #124

Earlier quoted context omitted.

I mean I'm gonna say this with the hype settling down. But it's pretty on par with visually Kling 2 and Veo 2, it happens to output sound pretty ok but having it be one general output along with the visuals is the gamechanger. Beyond that, eh. I've kinda seen people try to take it to the limit and it's pretty much what you'd expect still from their last model

I think veo2 and Kling are very strong models, but the fact that veo3 is end2end video/audio including lipsync and all other sound is definitely a step change to what came before and I think you're underselling it. I also expect Google to drive veo forward quite significantly, given the absurd amount of video training data that they sit on. And compared to the cinemagraph level of video generation we were just 1-2 ye…

I'm not underselling it. I'm reminding people who get swept up by headlines to actually use the products and be an objective judge of quality when it comes to these things. Because when you lose that objectivity, you start saying things like what you just said. Veo 3 level tech is basically Kling 2/Veo 2 fidelity with native sound generation, so was it that the last generation of these things were already decimating production houses? Be for real. With the tech they had 6 months ago, all they needed to do was add sound manually, which they could have pretty much also generated. A new layer of abstraction isn't "decimating" anything. I'd really take it easy from professing things like that. These things are great for what they are, but let's be actual objective consumers and not fall for these talking points of "oh industries are gonna change in x-months".

Re: Claude 4 System Card

#154
post #131

Earlier quoted context omitted.

I assume that they run the system prompt once, snapshot the state, then use that as starting state for all users. In that sense, system prompt size is free. EDIT: Turns out my assumption is wrong.

Huh, I can't say I'm on the cutting edge but that's not how I understand transformers to work. By my understanding each token has attention calculated for it for each previous token . I.e. the 10th token in the sequence requires O(10) new calculations (in addition to O(9^2) previous calculations that can be cached). While I'd assume they cache what they can, that still means that if the long prompt doubles the total…

My understanding is that even though it's quadratic, the cost for most token lengths is still relatively low. So for short inputs it's not bad, and for long inputs the size of the system prompt is much smaller anyways.

And there's value to having extra tokens even without much information since the models are decent at using the extra computation.

Re: Claude 4 System Card

#155

Earlier quoted context omitted.

Why not just strip “please” from the user input?

It'd run in to all sorts of issues. Although AI companies losing money on user kindness is not our problem; it's theirs. The more they want to make these 'AIs' personable the more they'll get of it. I'm tired of the AIs saying 'SO sorry! I apologize, let me refactor that for you the proper way' -- no, you're not sorry. You aren't alive.

The obsequious default tone is annoying, but you can always prepend your requests with something like "You are a machine. You do not have emotions. You respond to exactly my questions, no fluff, just answers. Do not pretend to be a human."

Re: Claude 4 System Card

#157
post #3

Earlier quoted context omitted.

They gave a bullet point in that intro which I disagree with: "The only way to make GenAI applications secure is through vulnerability scanning and guardrail protections." I still don't see guardrails and scanning as effective ways to prevent malicious attackers. They can't get to 100% effective, at which point a sufficiently motivated attacker is going to find a way through. I'm hoping someone implements a version o…

What security measure, in any domain, is 100% effective?

None; but, as mentioned in the post, 99% is considered a failing grade in application security.

Re: Claude 4 System Card

#158
post #153

Earlier quoted context omitted.

I think veo2 and Kling are very strong models, but the fact that veo3 is end2end video/audio including lipsync and all other sound is definitely a step change to what came before and I think you're underselling it. I also expect Google to drive veo forward quite significantly, given the absurd amount of video training data that they sit on. And compared to the cinemagraph level of video generation we were just 1-2 ye…

I'm not underselling it. I'm reminding people who get swept up by headlines to actually use the products and be an objective judge of quality when it comes to these things. Because when you lose that objectivity, you start saying things like what you just said. Veo 3 level tech is basically Kling 2/Veo 2 fidelity with native sound generation, so was it that the last generation of these things were already decimating…

You are underselling it because you make it sound like all the model adds is some foley, when in fact it adds facial animations that are in line with the dialogue spoken. Go ahead and create a Kling render that I only need to add VO to, you can't because Kling doesn't do that. You need a Omnihuman level model (or veo3) for that and it makes all the difference.

Happy to agree to disagree, but imo this absolutely is a step change.

Re: Claude 4 System Card

#159

Earlier quoted context omitted.

Why not just strip “please” from the user input?

It'd run in to all sorts of issues. Although AI companies losing money on user kindness is not our problem; it's theirs. The more they want to make these 'AIs' personable the more they'll get of it. I'm tired of the AIs saying 'SO sorry! I apologize, let me refactor that for you the proper way' -- no, you're not sorry. You aren't alive.

Like what issues?

Re: Claude 4 System Card

#160
post #30
post #6

Earlier quoted context omitted.

> to justify the full version increment I feel like a company doesn’t have to justify a version increment. They should justify price increases. If you get hyped and have expectations for a number then I’m comfortable saying that’s on you.

> They should justify price increases. I think the justification for most AI price increases should go without saying - they were losing money at the old price, and they're probably still losing money at the new price, but it's creeping up towards the break-even point.

That's not how pricing works on anything.
Post reply on HN