Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

831–840 of 1001 posts

Re: Claude Sonnet 4.6

#831

I always grew up hearing “competition is good for the consumer.” But I never really internalized how good fierce battles for market share are. The amount of competition in a space is directly proportional to how good the results are for consumers.

> how good the results are for consumers.

Only if you take consummer electronics out of the equation, because this AI arm race has wrecked havoc in the market for consumer GPUs, RAM, SSD and HDD.

If you take the arm race externalities into account, I'm very much unconvinced that we're better off than last year.

Re: Claude Sonnet 4.6

#832
post #192

The weirdest thing about this AI revolution is how smooth and continuous it is. If you look closely at differences between 4.6 and 4.5, it’s hard to see the subtle details. A year ago today, Sonnet 3.5 (new), was the newest model. A week later, Sonnet 3.7 would be released. Even 3.7 feels like ancient history! But in the gradient of 3.5 to 3.5 (new) to 3.7 to 4 to 4.1 to 4.5, I can’t think of one moment where I saw e…

If you've been using each new step is very noticeable and so have the mindshare. Around Sonnet 3.7 Claude Code-style coding became usable, and very quickly gained a lot of marketshare. Opus 4 could tackle significant more complexity. Opus 4.6 has been another noticable step up for me, suddenly I can let CC run significantly more independently, allowing multiple parallel agents where previously too much babysitting wa…

I think this is where there's a huge distinction between ability/performance/benchmark figures and utility. You can have smooth improvements to performance, but marked step changes in utility as they cross thresholds where you're able to use them for new tasks.

Re: Claude Sonnet 4.6

#833
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

Not for the entire world, with their pricing it is only good for US market, for the rest of the world we have ChatGPT and cheaper Chinese models.

Re: Claude Sonnet 4.6

#834

I always grew up hearing “competition is good for the consumer.” But I never really internalized how good fierce battles for market share are. The amount of competition in a space is directly proportional to how good the results are for consumers.

Remember when GPT-2 was “too dangerous to release” in 2019? That could have still been the state in 2026 if they didn’t YOLO it and ship ChatGPT to kick off this whole race.

I think the diffusion model race would've kicked off anyway. Didn't it even start before ChatGPT was released?

I think the spark would've been lit either way.

It's kind of funny how both of these things kicked off within a few months.

Re: Claude Sonnet 4.6

#835
post #525

It doesn't do so well on my stupid benchmarks, lol: https://aibenchy.com Gets wrong some tests. It does answer correctly, BUT it doesn't respect the request to respond ONLY with the answer, it keeps adding extra explanations at the end.

Looks like you're mixing up two things when testing: the correct answer and format following. If you want both, why not use https://platform.claude.com/docs/en/build-with-claude/struct... ? If you don't care about the structure, why penalise the correct answers? In realistic usage people don't say "I really care about the format a lot... but not enough to guarantee it".

Because the format can't also be strictly defined via structured output, and you have to write it in plain words. Imagine you also have a field within your JSON, which also needs a specific format. It's AI, you don't want to write a 2000lines JSON schema to define what you need and how to parse it, that's the point of using AI instead of writing your own data extraction script.

Also, simply because a human would respect it properly. And it's quite clear what the request was.

Thanks for the suggestion to separate format following from correct answer, good idea, I'll think about it.

Still, some good AIs do it properly, and as expectedly, why would I change the tests specifically for Claude, which is basically the only one with this problem.

Re: Claude Sonnet 4.6

#836
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

Their goal is to monopolize labor for anything that has to do with i/o on a computer, which is way more than SWE. Its simple, this technology literally cannot create new jobs it simply can cause one engineer (or any worker whos job has to do with computer i/o) to do the work of 3, therefore allowing you to replace workers (and overwork the ones you keep). Companies don't need "more work" half the "features"/"products…

I always find these "anti-AI" AI believer takes fascinating. If true AGI (which you are describing) comes to pass, there will certainly be massive societal consequences, and I'm not saying there won't be any dangers. But the economics in the resulting post-scarcity regime will be so far removed from our current world that I doubt any of this economic analysis will be even close to the mark.

I think the disconnect is that you are imagining a world where somehow LLMs are able to one-shot web businesses, but robotics and real-world tech is left untouched. Once LLMs can publish in top math/physics journals with little human assistance, it's a small step to dominating NeurIPS and getting us out of our mini-winter in robotics/RL. We're going to have Skynet or Star Trek, not the current weird situation where poor people can't afford healthy food, but can afford a smartphone.

Re: Claude Sonnet 4.6

#837

Earlier quoted context omitted.

Their goal is to monopolize labor for anything that has to do with i/o on a computer, which is way more than SWE. Its simple, this technology literally cannot create new jobs it simply can cause one engineer (or any worker whos job has to do with computer i/o) to do the work of 3, therefore allowing you to replace workers (and overwork the ones you keep). Companies don't need "more work" half the "features"/"products…

I always find these "anti-AI" AI believer takes fascinating. If true AGI (which you are describing) comes to pass, there will certainly be massive societal consequences, and I'm not saying there won't be any dangers. But the economics in the resulting post-scarcity regime will be so far removed from our current world that I doubt any of this economic analysis will be even close to the mark. I think the disconnect is…

> We're going to have Skynet or Star Trek

Star Trek only got a good society after an awful war, so neither of these options are good.

Re: Claude Sonnet 4.6

#838

> i need to wash my helicopter at the helicopter wash. it is 50m away, should i walk or fly there with my helicopter. Sonnet 4.6: Walk! Flying a helicopter 50 metres would be more trouble than it's worth — by the time you've done your pre-flight checks, spun up the rotors, lifted off, and then safely landed again, you'd have walked there and back twice. Just stroll over.

Asked gemini and it said to use ground handling wheels. I think it actually makes sense to use that for this distance.

Re: Claude Sonnet 4.6

#840
post #731

They use the word "Sonnet" 60+ times on that page but never give the casual reader any context of what a "Sonnet model" actually is. Neither does their landing page. You have to scroll all the way to the footer to find a link under the "Models" section. You click it and you finally get the description "Hybrid reasoning model with superior intelligence for agents, featuring a 1M context window" You then compare that t…

I won't argue with your point; both Anthropic and OpenAI name their models poorly, and it is hard to follow unless you're already following it. "Sonnet" only makes sense relative to other things but not by itself. If you don't know those other things, it is difficult to understand. But, if you were asking (and I'm not sure that you are): "Sonnet 4.6 is a cheaper, but worse, version of Opus 4.6 which itself is like GP…

> But, if you were asking (and I'm not sure that you are)

I was wondering, so thank you!

Post reply on HN