Has anyone tested how good the 1M context window is? i.e given an actual document, 1M tokens long. Can you ask it some question that relies on attending to 2 different parts of the context, and getting a good repsonse? I remember folks had problems like this with Gemini. I would be curious to see how Sonnet 4.6 stands up to it.
Did you see the graph benchmark? I found it quite interesting. It had to do a graph traversal on a natural text representation of a graph. Pretty much your problem.
Claude Sonnet 4.6
381–390 of 1001 posts
Re: Claude Sonnet 4.6
#382Does anyone know when will possibly arrive 1M context windows to at least MAX x20 subscriptions for claude code? I would even pay x50 if it allowed that. API usage is too expensive.
Re: Claude Sonnet 4.6
#383Earlier quoted context omitted.
The X grok feature is one of the best end user feature or large scale genai
What?! That's well regarded as one of the worst features introduced after the Twitter acquisition. Any thread these days is filled with "@grok is this true?" low effort comments. Not to mention the episode in which people spent two weeks using Grok to undress underage girls.
Re: Claude Sonnet 4.6
#384Earlier quoted context omitted.
The X grok feature is one of the best end user feature or large scale genai
That's news to me, I haven't read a single Grok post in my life. Am I missing out?
Re: Claude Sonnet 4.6
#385Earlier quoted context omitted.
> politically aligned AI company Like grok/xAI you mean?
I meant in a general sense. grok/xAI are politically aligned with whatever Musk wants. I haven't used their products but yes they're likely harmful in some ways. My concern is more over time if the federal government takes a more active role in trying to guide corporate behavior to align with moral or political goals. I think that's already occurring with the current administration but over a longer period of time if…
Re: Claude Sonnet 4.6
#386Earlier quoted context omitted.
We see the same with Google's Flash models. It's easier to make a small capable model when you have a large model to start from.
Flash models are nowhere near Pro models in daily use. Much higher hallucinations, and easy to get into a death sprawl of failed tool uses and never come out You should always take those claim that smaller models are as capable as larger models with a grain of salt.
Re: Claude Sonnet 4.6
#387Earlier quoted context omitted.
And you believe the other open source models are a signal for ethics? Don't have a dog in this fight, haven't done enough research to proclaim any LLM provider as ethical but I pretty much know the reason Meta has an open source model isn't because they're good guys.
> Don't have a dog in this fight, That's probably why you don't get it, then. Facebook was the primary contributor behind Pytorch, which basically set the stage for early GPT implementations. For all the issues you might have with Meta's social media, Facebook AI Research Labs have an excellent reputation in the industry and contributed greatly to where we are now. Same goes for Google Brain/DeepMind despite their Go…
Say I'm your neighbor and I make a move on your wife, your wife tells you this. Now I'm hosting a BBQ which is free for all to come, everyone in the neighborhood cheers for me. A neighbor praises me for helping him fix his car.
Someone asks you if you're coming to the BBQ, you say to him nah.. you don't like me. They go, 'WHAT? jack_pp? He rescues dogs and helped fix my roof! How can you not like him?'
Re: Claude Sonnet 4.6
#388I’m voting with my dollars by having cancelled my ChatGPT subscription and instead subscribing to Claude. Google needs stiff competition and OpenAI isn’t the camp I’m willing to trust. Neither is Grok. I’m glad Anthropic’s work is at the forefront and they appear, at least in my estimation, to have the strongest ethics.
This is just you verifying that their branding is working. It signals nothing about their actual ethics.
Re: Claude Sonnet 4.6
#389Many people have reported Opus 4.6 is a step back from Opus 4.5 - that 4.6 is consuming 5-10x as many tokens as 4.5 to accomplish the same task: https://github.com/anthropics/claude-code/issues/23706 I haven't seen a response from the Anthropic team about it. I can't help but look at Sonnet 4.6 in the same light, and want to stick with 4.5 across the board until this issue is acknowledged and resolved.
Those suggest opposite things about anthropic’s profit margins.
I’m not convinced 4.6 is much better than 4.5. The big discontinuous breakthroughs seem to be due to how my code and tests are structured, not model bumps.
Re: Claude Sonnet 4.6
#390I always grew up hearing “competition is good for the consumer.” But I never really internalized how good fierce battles for market share are. The amount of competition in a space is directly proportional to how good the results are for consumers.