Claude 4
anthropic.com
Claude 4
1–10 of 1001 posts
Re: Claude 4
#2my highlights:
1. Coding ability: "Claude Opus 4 is our most powerful model yet and the best coding model in the world, leading on SWE-bench (72.5%) and Terminal-bench (43.2%). It delivers sustained performance on long-running tasks that require focused effort and thousands of steps, with the ability to work continuously for several hours—dramatically outperforming all Sonnet models and significantly expanding what AI agents can accomplish." however this is Best of N, with no transparency on size of N and how they decide the best, saying "We then use an internal scoring model to select the best candidate from the remaining attempts." Claude Code is now generally available (we covered in http://latent.space/p/claude-code )
2. Memory highlight: "Claude Opus 4 also dramatically outperforms all previous models on memory capabilities. When developers build applications that provide Claude local file access, Opus 4 becomes skilled at creating and maintaining 'memory files' to store key information. This unlocks better long-term task awareness, coherence, and performance on agent tasks—like Opus 4 creating a 'Navigation Guide' while playing Pokémon." Memory Cookbook: https://github.com/anthropics/anthropic-cookbook/blob/main/t...
3. Raw CoT available: "we've introduced thinking summaries for Claude 4 models that use a smaller model to condense lengthy thought processes. This summarization is only needed about 5% of the time—most thought processes are short enough to display in full. Users requiring raw chains of thought for advanced prompt engineering can contact sales about our new Developer Mode to retain full access."
4. haha: "We no longer include the third ‘planning tool’ used by Claude 3.7 Sonnet. " 5. context caching now has a premium 1hr TTL option: "Developers can now choose between our standard 5-minute time to live (TTL) for prompt caching or opt for an extended 1-hour TTL at an additional cost"
6. https://www.anthropic.com/news/agent-capabilities-api new code execution tool (sandbox) and file tool
Re: Claude 4
#3Re: Claude 4
#4Quoted post unavailable.
Have you used it?
I liked Claude 3.7 but without context this comes off as what the kids would call "glazing"
Re: Claude 4
#5Edit: How do you install it? Running `/ide` says "Make sure your IDE has the Claude Code extension", where do you get that?
Re: Claude 4
#6ETA: I guess Anthropic still thinks they can command a premium, I hope they're right (because I would love to pay more for smarter models).
> Pricing remains consistent with previous Opus and Sonnet models: Opus 4 at $15/$75 per million tokens (input/output) and Sonnet 4 at $3/$15.
Re: Claude 4
#7Re: Claude 4
#8livestream here: https://youtu.be/EvtPBaaykdo my highlights: 1. Coding ability: "Claude Opus 4 is our most powerful model yet and the best coding model in the world, leading on SWE-bench (72.5%) and Terminal-bench (43.2%). It delivers sustained performance on long-running tasks that require focused effort and thousands of steps, with the ability to work continuously for several hours—dramatically outperforming all So…
Re: Claude 4
#9Re: Claude 4
#10Quoted post unavailable.