Live data from Hacker News

Claude Opus 4.6

anthropic.com

161–170 of 1001 posts

Re: Claude Opus 4.6

#161
post #62

Claude Code release notes: > Version 2.1.32: • Claude Opus 4.6 is now available! • Added research preview agent teams feature for multi-agent collaboration (token-intensive feature, requires setting CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1) • Claude now automatically records and recalls memories as it works • Added "Summarize from here" to the message selector, allowing partial conversation summarization. • Skills defi…

> Claude now automatically records and recalls memories as it works Neat: https://code.claude.com/docs/en/memory I guess it's kind of like Google Antigravity's "Knowledge" artifacts?

Are we sure the docs page has been updated yet? Because that page doesn't say anything about automatic recording of memories.

Re: Claude Opus 4.6

#162
post #46

The benchmarks are cool and all but 1M context on an Opus-class model is the real headline here imo. Has anyone actually pushed it to the limit yet? Long context has historically been one of those "works great in the demo" situations.

Paying $10 per request doesn't have me jumping at the opportunity to try it!

Re: Claude Opus 4.6

#163

5.3 codex https://openai.com/index/introducing-gpt-5-3-codex/ crushes with a 77.3% in Terminal Bench. The shortest lived lead in less than 35 minutes. What a time to be alive!

claude swe-bench is 80.8 and codex is 56.8

Seems like 4.6 is still all-around better?

Re: Claude Opus 4.6

#167

5.3 codex https://openai.com/index/introducing-gpt-5-3-codex/ crushes with a 77.3% in Terminal Bench. The shortest lived lead in less than 35 minutes. What a time to be alive!

claude swe-bench is 80.8 and codex is 56.8 Seems like 4.6 is still all-around better?

Its SWE bench pro not swe bench verified. The verified benchmark has stagnated

Re: Claude Opus 4.6

#169

5.3 codex https://openai.com/index/introducing-gpt-5-3-codex/ crushes with a 77.3% in Terminal Bench. The shortest lived lead in less than 35 minutes. What a time to be alive!

The lack of broad benchmark reports in this makes me curious: Has OpenAI reverted to benchmaxxing? Looking forward to hearing opinions once we all try both of these out

Re: Claude Opus 4.6

#170
post #40

The bicycle frame is a bit wonky but the pelican itself is great: https://gist.github.com/simonw/a6806ce41b4c721e240a4548ecdbe...

There's no way they actually work on training this.

I suspect they're training on this.

I asked Opus 4.6 for a pelican riding a recumbent bicycle and got this.

https://i.imgur.com/UvlEBs8.png

Post reply on HN