Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

381–390 of 1001 posts

Re: Claude Sonnet 4.6

#381

Has anyone tested how good the 1M context window is? i.e given an actual document, 1M tokens long. Can you ask it some question that relies on attending to 2 different parts of the context, and getting a good repsonse? I remember folks had problems like this with Gemini. I would be curious to see how Sonnet 4.6 stands up to it.

Did you see the graph benchmark? I found it quite interesting. It had to do a graph traversal on a natural text representation of a graph. Pretty much your problem.

Update: I took a corpus of personal chat data (this way it wouldn't be seen in training), and tried asking it some paraphrased questions. It performed quite poorly.

Re: Claude Sonnet 4.6

#382

Does anyone know when will possibly arrive 1M context windows to at least MAX x20 subscriptions for claude code? I would even pay x50 if it allowed that. API usage is too expensive.

I don't know when it will be included as part of the subscription in Claude Code, but at least it's a paid add-on in the MAX plan now. That's a decent alternative for situations where the extra space is valuable, especially without having to setup/maintain API billing separately.

Re: Claude Sonnet 4.6

#383

Earlier quoted context omitted.

The X grok feature is one of the best end user feature or large scale genai

What?! That's well regarded as one of the worst features introduced after the Twitter acquisition. Any thread these days is filled with "@grok is this true?" low effort comments. Not to mention the episode in which people spent two weeks using Grok to undress underage girls.

high adoption means this works...

Re: Claude Sonnet 4.6

#384

Earlier quoted context omitted.

The X grok feature is one of the best end user feature or large scale genai

That's news to me, I haven't read a single Grok post in my life. Am I missing out?

im talking about the "explain this post" feature on top right of a message where groks mix thread data, live data and other tweets to unify a stream of information

Re: Claude Sonnet 4.6

#385
post #246

Earlier quoted context omitted.

> politically aligned AI company Like grok/xAI you mean?

I meant in a general sense. grok/xAI are politically aligned with whatever Musk wants. I haven't used their products but yes they're likely harmful in some ways. My concern is more over time if the federal government takes a more active role in trying to guide corporate behavior to align with moral or political goals. I think that's already occurring with the current administration but over a longer period of time if…

I don’t think people will just accept that. They‘ll use some European or Chinese model instead that doesn’t have that problem.

Re: Claude Sonnet 4.6

#386

Earlier quoted context omitted.

We see the same with Google's Flash models. It's easier to make a small capable model when you have a large model to start from.

Flash models are nowhere near Pro models in daily use. Much higher hallucinations, and easy to get into a death sprawl of failed tool uses and never come out You should always take those claim that smaller models are as capable as larger models with a grain of salt.

Flash model n is generally a slightly better Pro model (n-1), in other words you get to use the previously premium model as a cheaper/faster version. That has value.

Re: Claude Sonnet 4.6

#387

Earlier quoted context omitted.

And you believe the other open source models are a signal for ethics? Don't have a dog in this fight, haven't done enough research to proclaim any LLM provider as ethical but I pretty much know the reason Meta has an open source model isn't because they're good guys.

> Don't have a dog in this fight, That's probably why you don't get it, then. Facebook was the primary contributor behind Pytorch, which basically set the stage for early GPT implementations. For all the issues you might have with Meta's social media, Facebook AI Research Labs have an excellent reputation in the industry and contributed greatly to where we are now. Same goes for Google Brain/DeepMind despite their Go…

A hired assassin can have an excellent reputation too. What does that have to do with ethics?

Say I'm your neighbor and I make a move on your wife, your wife tells you this. Now I'm hosting a BBQ which is free for all to come, everyone in the neighborhood cheers for me. A neighbor praises me for helping him fix his car.

Someone asks you if you're coming to the BBQ, you say to him nah.. you don't like me. They go, 'WHAT? jack_pp? He rescues dogs and helped fix my roof! How can you not like him?'

Re: Claude Sonnet 4.6

#388

I’m voting with my dollars by having cancelled my ChatGPT subscription and instead subscribing to Claude. Google needs stiff competition and OpenAI isn’t the camp I’m willing to trust. Neither is Grok. I’m glad Anthropic’s work is at the forefront and they appear, at least in my estimation, to have the strongest ethics.

This is just you verifying that their branding is working. It signals nothing about their actual ethics.

Unfortunately, you're correct. Claude was used in the Venezuela raid, Anthropic's consent be damned. They're not resisting, they're marketing resistence.

Re: Claude Sonnet 4.6

#389

Many people have reported Opus 4.6 is a step back from Opus 4.5 - that 4.6 is consuming 5-10x as many tokens as 4.5 to accomplish the same task: https://github.com/anthropics/claude-code/issues/23706 I haven't seen a response from the Anthropic team about it. I can't help but look at Sonnet 4.6 in the same light, and want to stick with 4.5 across the board until this issue is acknowledged and resolved.

I’ve noticed the opaque weekly quota meter goes up more slowly with 4.6, but it more frequently goes off and works for an hour+, with really high reported token counts.

Those suggest opposite things about anthropic’s profit margins.

I’m not convinced 4.6 is much better than 4.5. The big discontinuous breakthroughs seem to be due to how my code and tests are structured, not model bumps.

Re: Claude Sonnet 4.6

#390

I always grew up hearing “competition is good for the consumer.” But I never really internalized how good fierce battles for market share are. The amount of competition in a space is directly proportional to how good the results are for consumers.

I grew up with every service enshitified in the end. Whoever has more money wins the race and gets richer, that's free market for ya.
Post reply on HN