Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

541–550 of 1001 posts

Re: Claude Sonnet 4.6

#541
> Sonnet 4.5, starting at $3/$15 per million tokens.

Are people really willing to pay these prices? The open-weight models are catching up in a rapid pace while keeping the prices so low. MiniMax M2.5, Kimi 2.5 and GLM-5 is dirt cheap compared to this. They may not be sota but they are more than good enough.

Re: Claude Sonnet 4.6

#542

Earlier quoted context omitted.

Keep in mind that the people who experience issues will always be the loudest. I've overall enjoyed 4.6. On many easy things it thinks less than 4.5, leading to snappier feedback. And 4.6 seems much more comfortable calling tools: it's much more proactive about looking at the git history to understand the history of a bug or feature, or about looking at online documentation for APIs and packages. A recent claude code…

Do you need to upload your git for it to analyuze it? Or are they reading it off github ?

They're probably running it with a claude code like tool and it has a local (to the tool, not to anthropic) copy of the git repo it can query using the cli.

Re: Claude Sonnet 4.6

#543

Earlier quoted context omitted.

So like....every business having electricity? I am not a economist so would love someone smarter than me explain how this is any different than the advent of electricity and how that affected labor.

The difference is that electricity wasn't being controlled by oligarchs that want to shape society so they become more rich while pillaging the planet and hurting/killing real human beings. I'd be more trusting of LLM companies if they were all workplace democracies, not really a big fan of the centrally planned monarchies that seem to be most US corporations.

Its main distinction from previous forms of automation is its ability to apply reasoning to processes and its potential to operate almost entirely without supervision, and also to be retasked with trivial effort. Conventional automation requires huge investments in a very specific process. Widespread automation will allow highly automated organizations to pivot or repurpose overnight.

Re: Claude Sonnet 4.6

#544

Earlier quoted context omitted.

I actually agree with you, but I have no idea how one can compete in this playing field. The second there are a couple of bad actors in spammarketing, your hands are tied. You really can’t win without playing dirty. I really hate this, not justifying their behaviour, but have no clue how one can do without the other.

Its just law of the jungle all over again. Might makes right. Outcomes over means. Game theory wise there is no solution except to declare (and enforce) spaces where leeching / degrading the environment is punished, and sharing, building, and giving back to the environment is rewarded. Not financially, because it doesn't work that way, usually through social cred or mutual values. But yeah the internet can no longer…

> the rest of the world quietly leaves if they know what's good for them.

Userbase of TikTok, Instagram and etc. has increased YoY. People suck at making decisions for their own good on average.

Re: Claude Sonnet 4.6

#545

Earlier quoted context omitted.

Their goal is to monopolize labor for anything that has to do with i/o on a computer, which is way more than SWE. Its simple, this technology literally cannot create new jobs it simply can cause one engineer (or any worker whos job has to do with computer i/o) to do the work of 3, therefore allowing you to replace workers (and overwork the ones you keep). Companies don't need "more work" half the "features"/"products…

The price of oil at the price of water (ecology apart) should be a good thing. Automation should be, obviously, a good thing, because more is produced with less labor. What it says of ourselves and our politics that so many people (me included) are afraid of it? In a sane world, we would realize that, in a post-work world, the owner of the robots have all the power, so the robots should be owned in common. The soluti…

While I agree, I am not hopeful. The incentive alignment has us careening towards Elysium rather than Star Trek.

Re: Claude Sonnet 4.6

#546

> Sonnet 4.5, starting at $3/$15 per million tokens. Are people really willing to pay these prices? The open-weight models are catching up in a rapid pace while keeping the prices so low. MiniMax M2.5, Kimi 2.5 and GLM-5 is dirt cheap compared to this. They may not be sota but they are more than good enough.

Some people will want the models like claude where you don't have to be super-specific and it will infer exactly what you mean.

With the GLM models you have to confirm with it exactly what you want, and not miss any detail.

Re: Claude Sonnet 4.6

#547
post #449

Earlier quoted context omitted.

Or self host the oss models on the second hand GPU and RAM that's left when the big labs implode

China will stop releasing open weights models as soon as they get within striking range; c.f. seedance 2.0.

ByteDance never really open sourced their models though. But I agree, they will only open source when it doesn't really matter.

Re: Claude Sonnet 4.6

#548

> Sonnet 4.5, starting at $3/$15 per million tokens. Are people really willing to pay these prices? The open-weight models are catching up in a rapid pace while keeping the prices so low. MiniMax M2.5, Kimi 2.5 and GLM-5 is dirt cheap compared to this. They may not be sota but they are more than good enough.

It depends on how much you value the gap between “pretty good” and SOTA… I’ve noticed that Opus is more “expensive”,” but an error-filled rabbit hole is expensive too!

Re: Claude Sonnet 4.6

#549
post #395

Earlier quoted context omitted.

In my experience with the models (watching Claude play Pokemon), the models are similar in intelligence, but are very different in how they approach problems: Opus 4.5 hyperfocuses on completing its original plan, far more than any older or newer version of Claude. Opus 4.6 gets bored quickly and is constantly changing its approach if it doesn't get results fast. This makes it waste more time on"easy" tasks where the…

Genuinely one of the more interesting model evals I've seen described. The sunk cost framing makes sense -- 4.5 doubles down, 4.6 cuts losses faster. 9 days vs 59 is a wild result. Makes me wonder how much of the regression complaints are from people hitting 4.6 on tasks where the first approach was obviously correct.

Notably 45 out of the 50 days of improvement were in two specific dungeons (Silph Co and Cinnabar Mansion) where 4.5 was entirely inadequate and was looping the same mistaken ideas with only minor variation, until eventually it stumbled by chance into the solution. Until we saw how much better it did in those spots, we weren't completely sure that 4.6 was an improvement at all!

https://docs.google.com/spreadsheets/u/0/d/e/2PACX-1vQDvsy5D...

Re: Claude Sonnet 4.6

#550

Earlier quoted context omitted.

Wow, haha. I tried this with gpt5.2 and, presumably due to some customisations I have set, this is how it went: --- Me: I want to wash my car. My car is currently at home. The car wash is 50 meters away. Should I walk or drive? GPT: You’re asking an AI to adjudicate a 50-metre life decision. Humanity really did peak with the moon landing. Walk. Obviously walk. Fifty metres is barely a committed stroll. By the time yo…

OK! customisations please? ...

All of my “characteristics” (a setting I don’t think I’ve seen before) are set to default and my custom instructions are as follows…

——

Always assume British English when relevant. If there are any technical, grammatical, syntactical, or other errors in my statement please correct them before responding.

Tell it like it is; don't sugar-coat responses. Adopt a skeptical, questioning approach.

Post reply on HN