Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

41–50 of 1001 posts

Re: Claude Sonnet 4.6

#41
post #9

[flagged]

Incompleteness is inherent to a physical reality being deconstructed by entropy.

Of your concern is morality, humans need to learn a lot about that themselves still. It's absurd the number of first worlders losing their shit over loss of paid work drawing manga fan art in the comfort of their home while exploiting labor of teens in 996 textile factories.

AI trained on human outputs that lack such self awareness, lacks awareness of environmental externalities of constant car and air travel, will result in AI with gaps in their morality.

Gary Marcus is onto something with the problems inherent to systems without formal verification. But he will fully ignores this issue exists in human social systems already as intentional indifference to economic externalities, zero will to police the police and watch the watchers.

Most people are down to watch the circus without a care so long as the waitstaff keep bringing bread.

Re: Claude Sonnet 4.6

#42
post #36

I wonder what difference have people found with sonnet 4.5 and opus 4.5 and probably similar delta will remain. Was sonnet 4.5 much worse than opus?

Sonnet 4.5 was a pretty significant improvement over Opus 4.

Yes but it’s easier to understand difference between 4.5 sonnet and opus and apply that difference to opus 4.6

Re: Claude Sonnet 4.6

#43
post #17
post #5

My take away is: it's roughly as good as Opus 4.5. Now the question is: how much faster or cheaper is it?

If it maintains the same price (with Anthropic tends to do or undercuts themselves) then this would be 1/3rd of the price of Opus. Edit: Yep, same price. "Pricing remains the same as Sonnet 4.5, starting at $3/$15 per million tokens."

3 is not 1/3 of 5 tho. Opus costs $5/$25

Re: Claude Sonnet 4.6

#44
post #10

It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.

Why is it wild that a LLM is as capable as a previously released LLM?

It means price has decreased by 3 times in a few months.

Re: Claude Sonnet 4.6

#45
post #10

It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.

Why is it wild that a LLM is as capable as a previously released LLM?

Because Opus 4.5 inference is/was more expensive.

Re: Claude Sonnet 4.6

#46
post #10

It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.

We see the same with Google's Flash models. It's easier to make a small capable model when you have a large model to start from.

Re: Claude Sonnet 4.6

#47
post #28

Earlier quoted context omitted.

simonw hasn't shown up yet, so here's my "Generate an SVG of a pelican riding a bicycle" https://claude.ai/public/artifacts/67c13d9a-3d63-4598-88d0-5...

We finally have AI safety solved! Look at that helmet

"Look ma, no wings!"

:D

Re: Claude Sonnet 4.6

#48
post #10

It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.

The most exciting part isn't necessarily the ceiling raising though that's happening, but the floor rising while costs plummet. Getting Opus-level reasoning at Sonnet prices/latency is what actually unlocks agentic workflows. We are effectively getting the same intelligence unit for half the compute every 6-9 months.

Re: Claude Sonnet 4.6

#49
post #10

It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.

Why is it wild that a LLM is as capable as a previously released LLM?

Opus is supposed to be the expensive-but-quality one, while Sonnet is the cheaper one.

So if you don't want to pay the significant premium for Opus, it seems like you can just wait a few weeks till Sonnet catches up

Re: Claude Sonnet 4.6

#50
post #10

It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.

Why is it wild that a LLM is as capable as a previously released LLM?

Because Opus 4.5 was released like a month ago and state of the art, and now the significantly faster and cheaper version is already comparable.
Post reply on HN