Live data from Hacker News

Claude 4

anthropic.com

131–140 of 1001 posts

Re: Claude 4

#131
post #83

Have they documented the context window changes for Claude 4 anywhere? My (barely informed) understanding was one of the reasons Gemini 2.5 has been so useful is that it can handle huge amounts of context --- 50-70kloc?

Context window is unchanged for Sonnet. (200k in/64k out): https://docs.anthropic.com/en/docs/about-claude/models/overv... In practice, the 1M context of Gemini 2.5 isn't that much of a differentiator because larger context has diminishing returns on adherence to later tokens.

Gemini's real strength is that it can stay on the ball even at 100k tokens in context.

Re: Claude 4

#132
post #37

Earlier quoted context omitted.

Gemini can beat the game?

Gemini has beat it already, but using a different and notably more helpful harness. The creator has said they think harness design is the most important factor right now, and that the results don't mean much for comparing Claude to Gemini.

Way offtopic to TFA now, but isn't using an improved harness a bit like saying "I'm going to hardcore as many priors as possible into this thing so it succeeds regardless of its ability to strategize, plan and execute?

Re: Claude 4

#133

Is this really worthy of a claude 4 label? Was there a new pre-training run? Cause this feels like 3.8... only swe went up significantly, and that as we all understand by now is done by cramming on specific post training data and doesn't generalize to intelligence. The agentic tooluse didn't improve and this says to me that it's not really smarter.

Slight decrease from Sonnet 3.7 in a few areas even. As always benchmarks say one thing, will need some practice with it to get a subjective opinion.

Re: Claude 4

#134
post #51

I've found myself having brand loyalty to Claude. I don't really trust any of the other models with coding, the only one I even let close to my work is Claude. And this is after trying most of them. Looking forward to trying 4.

I also recommend trying out Gemini, I'm really impressed by the latest 2.5. Let's see if Claude 4 makes me switch back.

Re: Claude 4

#136
From the system card [0]:

Claude Opus 4 - Knowledge Cutoff: Mar 2025 - Core Capabilities: Hybrid reasoning, visual analysis, computer use (agentic), tool use, adv. coding (autonomous), enhanced tool use & agentic workflows. - Thinking Mode: Std & "Extended Thinking Mode" Safety/Agency: ASL-3 (precautionary); higher initiative/agency than prev. models. 0/4 researchers believed that Claude Opus 4 could completely automate the work of a junior ML researcher.

Claude Sonnet 4 - Knowledge Cutoff: Mar 2025 - Core Capabilities: Hybrid reasoning - Thinking Mode: Std & "Extended Thinking Mode" - Safety: ASL-2.

[0] https://www-cdn.anthropic.com/4263b940cabb546aa0e3283f35b686...

Re: Claude 4

#137

Is this really worthy of a claude 4 label? Was there a new pre-training run? Cause this feels like 3.8... only swe went up significantly, and that as we all understand by now is done by cramming on specific post training data and doesn't generalize to intelligence. The agentic tooluse didn't improve and this says to me that it's not really smarter.

Benchmarks don't tell you as much as the actual coding vibe though

Re: Claude 4

#138
post #81
post #67

It’s been hard to keep up with the evolution in LLMs. SOTA models basically change every other week, and each of them has its own quirks. Differences in features, personality, output formatting, UI, safety filters… make it nearly impossible to migrate workflows between distinct LLMs. Even models of the same family exhibit strikingly different behaviors in response to the same prompt. Still, having to find each model’…

My advice: don't jump around between LLMs for a given project. The AI space is progressing too rapidly right now. Save yourself the sanity.

Each model has their own strength and weaknesses tho. You really shouldn’t be using one model for everything. Like, Claude is great at coding but is expensive so you wouldn’t use them for debugging to writing test benches. But the OpenAI models suck at architecture but are cheap, so are ideal for test benches, for example.

Re: Claude 4

#139
post #51

I've found myself having brand loyalty to Claude. I don't really trust any of the other models with coding, the only one I even let close to my work is Claude. And this is after trying most of them. Looking forward to trying 4.

i was the same, but then slowly converted to Gemini. Still not sure how that happened tbh
Post reply on HN