Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

441–450 of 1001 posts

Re: Claude Sonnet 4.6

#441
post #423

Earlier quoted context omitted.

I was just thinking earlier today how in an alternate universe, probably not too far removed from our own, Google has a monopoly on transformers and we are all stuck with a single GPT-3.5 level model, and Google has a GPT-4o model behind the scenes that it is terrified to release (but using heavily internally).

Now think about how often the patent system has stifled and stalled and delayed advancement for decades per innovation at a time. Where would we be if patents never existed?

Who knows? If we’d never moved on from trade secrets to patents, we might be a hundred years behind.

Re: Claude Sonnet 4.6

#442
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

If the world becomes dependent on computer-use than the AI buildout will be more than validated. That will require all that compute.

It will be validated but that doesn’t mean that the providers of these services will be making money. It’s about the demand at a profitable price. The uncontroversial part is that the demand exists at an unprofitable price.

Re: Claude Sonnet 4.6

#443

Many people have reported Opus 4.6 is a step back from Opus 4.5 - that 4.6 is consuming 5-10x as many tokens as 4.5 to accomplish the same task: https://github.com/anthropics/claude-code/issues/23706 I haven't seen a response from the Anthropic team about it. I can't help but look at Sonnet 4.6 in the same light, and want to stick with 4.5 across the board until this issue is acknowledged and resolved.

Keep in mind that the people who experience issues will always be the loudest. I've overall enjoyed 4.6. On many easy things it thinks less than 4.5, leading to snappier feedback. And 4.6 seems much more comfortable calling tools: it's much more proactive about looking at the git history to understand the history of a bug or feature, or about looking at online documentation for APIs and packages. A recent claude code…

Do you need to upload your git for it to analyuze it? Or are they reading it off github ?

Re: Claude Sonnet 4.6

#444

Earlier quoted context omitted.

Isn't "computer use" just interaction with a shell-like environment, which is routine for current agents?

Interesting question! In this context, "computer use" means the model is manipulating a full graphical interface, using a virtual mouse and keyboard to interact with applications (like Chrome or LibreOffice), rather than simply operating in a shell environment.

Indeed GUI-use would have been the better naming.

Re: Claude Sonnet 4.6

#445

I’m voting with my dollars by having cancelled my ChatGPT subscription and instead subscribing to Claude. Google needs stiff competition and OpenAI isn’t the camp I’m willing to trust. Neither is Grok. I’m glad Anthropic’s work is at the forefront and they appear, at least in my estimation, to have the strongest ethics.

I dropped ChatGPT as soon as they went to an ad supported model. Claude Opus 4.6 seems noticeably better than GPT 5.2 Thinking so far.

Re: Claude Sonnet 4.6

#446
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

Does it matter? Really? I can type awful stuff into a word processor. That's my fault, not the programs. So if I can trick an LLM into saying awful stuff, whose fault is that? It is also just a tool...

I can kill someone with a rock, a knife, a pistol, and a fully automatic rifle. There is a real difference in the other uses, efficacy, and scope of each.

Re: Claude Sonnet 4.6

#447

Many people have reported Opus 4.6 is a step back from Opus 4.5 - that 4.6 is consuming 5-10x as many tokens as 4.5 to accomplish the same task: https://github.com/anthropics/claude-code/issues/23706 I haven't seen a response from the Anthropic team about it. I can't help but look at Sonnet 4.6 in the same light, and want to stick with 4.5 across the board until this issue is acknowledged and resolved.

I wonder if it's actually from CC harness updates that make it much more inclined to use subagents, rather than from the model update.

Re: Claude Sonnet 4.6

#448
Still fails the car wash question, I took the prompt from the title of this thread: https://news.ycombinator.com/item?id=47031580

The answer was "Walk! It would be a bit counterproductive to drive a dirty car 50 meters just to get it washed — you'd barely move before arriving. Walking takes less than a minute, and you can simply drive it through the wash and walk back home afterward."

I've tried several other variants of this question and I got similar failures.

Re: Claude Sonnet 4.6

#449
post #165

Earlier quoted context omitted.

Until 2 remain, then it's extraction time.

Or self host the oss models on the second hand GPU and RAM that's left when the big labs implode

China will stop releasing open weights models as soon as they get within striking range; c.f. seedance 2.0.

Re: Claude Sonnet 4.6

#450

Still fails the car wash question, I took the prompt from the title of this thread: https://news.ycombinator.com/item?id=47031580 The answer was "Walk! It would be a bit counterproductive to drive a dirty car 50 meters just to get it washed — you'd barely move before arriving. Walking takes less than a minute, and you can simply drive it through the wash and walk back home afterward." I've tried several other variant…

Remarkable, since the goal is clearly stated and the language isn’t tricky.
Post reply on HN