Has anyone tested how good the 1M context window is? i.e given an actual document, 1M tokens long. Can you ask it some question that relies on attending to 2 different parts of the context, and getting a good repsonse? I remember folks had problems like this with Gemini. I would be curious to see how Sonnet 4.6 stands up to it.
Did you see the graph benchmark? I found it quite interesting. It had to do a graph traversal on a natural text representation of a graph. Pretty much your problem.
Claude Sonnet 4.6
61–70 of 1001 posts
Re: Claude Sonnet 4.6
#62Re: Claude Sonnet 4.6
#63Hoe much power did it take to train the models?
I would honestly guess that this is just a small amount of tweaking on top of the Sonnet 4.x models. It seems like providers are rarely training new 'base' models anymore. We're at a point where the gains are more from modifying the model's architecture and doing a "post" training refinement. That's what we've been seeing for the past 12-18 months, iirc.
Re: Claude Sonnet 4.6
#64With such a huge leap, i’m confused why they didn’t call it Sonnet 5? As someone who uses Sonnet 4.5 for 95% tasks due to costs, i’m pretty excited to try 4.6 at the same price
Re: Claude Sonnet 4.6
#65[flagged]
>we're just teaching them how to pass a polygraph. I understand the metaphor, but using 'pass a polygraph' as a measure of truthfulness or deception is dangerous in that it alludes to the polygraph as being a realistic measure of those metrics -- it is not.
A poly is only testing one thing: can you convince the polygrapher that you can lie successfully
Re: Claude Sonnet 4.6
#66It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.
Given that users prefered it to Sonnet 4.5 "only" in 70% of the cases (according to their blog post) makes me highly doubt that this is representative of real-life usage. Benchmarks are just completely meaningless.
Re: Claude Sonnet 4.6
#67It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.
We see the same with Google's Flash models. It's easier to make a small capable model when you have a large model to start from.
You should always take those claim that smaller models are as capable as larger models with a grain of salt.
Re: Claude Sonnet 4.6
#68[flagged]
It always has been. We already hit the point a while ag where we regularly caught them trying to be deceptive, so we should automatically assume from that point forward that if we don't catch them being deceptive, that may mean they're better at it rather than that they're not doing it.
Re: Claude Sonnet 4.6
#69A year ago today, Sonnet 3.5 (new), was the newest model. A week later, Sonnet 3.7 would be released.
Even 3.7 feels like ancient history! But in the gradient of 3.5 to 3.5 (new) to 3.7 to 4 to 4.1 to 4.5, I can’t think of one moment where I saw everything change. Even with all the noise in the headlines, it’s still been a silent revolution.
Am I just a believer in an emperor with no clothes? Or, somehow, against all probability and plausibility, are we all still early?
Re: Claude Sonnet 4.6
#70Earlier quoted context omitted.
> The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s. Yeah, but RAM prices are also back to 1990s levels.
Relief for you is available: https://computeradsfromthepast.substack.com/p/connectix-ram-...