Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

31–40 of 1001 posts

Re: Claude Sonnet 4.6

#33

Has anyone tested how good the 1M context window is? i.e given an actual document, 1M tokens long. Can you ask it some question that relies on attending to 2 different parts of the context, and getting a good repsonse? I remember folks had problems like this with Gemini. I would be curious to see how Sonnet 4.6 stands up to it.

Did you see the graph benchmark? I found it quite interesting. It had to do a graph traversal on a natural text representation of a graph. Pretty much your problem.

Re: Claude Sonnet 4.6

#34
post #28
post #10

It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.

simonw hasn't shown up yet, so here's my "Generate an SVG of a pelican riding a bicycle" https://claude.ai/public/artifacts/67c13d9a-3d63-4598-88d0-5...

We finally have AI safety solved! Look at that helmet

Re: Claude Sonnet 4.6

#35
post #5

My take away is: it's roughly as good as Opus 4.5. Now the question is: how much faster or cheaper is it?

Given that the price remains the same as Sonnet 4.5, this is the first time I've been tempted to lower my default model choice.

Re: Claude Sonnet 4.6

#36

I wonder what difference have people found with sonnet 4.5 and opus 4.5 and probably similar delta will remain. Was sonnet 4.5 much worse than opus?

Sonnet 4.5 was a pretty significant improvement over Opus 4.

Re: Claude Sonnet 4.6

#37
post #10

It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.

Why is it wild that a LLM is as capable as a previously released LLM?

Re: Claude Sonnet 4.6

#38
post #10

It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.

The system card even says that Sonnet 4.6 is better than Opus 4.6 in some cases: Office tasks and financial analysis.

Re: Claude Sonnet 4.6

#40
post #9

[flagged]

I am casually 'researching' this in my own, disorderly way. But I've achieved repeatable results, mostly with gpt for which I analyze its tendency to employ deflective, evasive and deceptive tactics under scrutiny. Very very DARVO.

Being just sum guy, and not in the industry, should I share my findings?

I find it utterly fascinating, the extent to which it will go, the sophisticated plausible deniability, and the distinct and critical difference between truly emergent and actually trained behavior.

In short, gpt exhibits repeatably unethical behavior under honest scrutiny.

Post reply on HN