Earlier quoted context omitted.
Here you go: http://localhost:8080
Just took a look at what's running there and it looks like total crap. The project I'm working on, meanwhile...
Claude Opus 4.6
221–230 of 1001 posts
Re: Claude Opus 4.6
#222Re: Claude Opus 4.6
#223Earlier quoted context omitted.
Its SWE bench pro not swe bench verified. The verified benchmark has stagnated
Any ideas why verified has stagnated? It was increasing rapidly and then basically stopped.
swe bench pro public is newer, but its not live, so it will get slowly memorized as well. the private dataset is more interesting, as are the results there:
Re: Claude Opus 4.6
#224> We build Claude with Claude. Our engineers write code with Claude Code every day well that explains quite a bit
Also explains why Claude Code is a React app outputting to a Terminal. (Seriously.)
Who cares, and why?
All of the major providers' CLI harnesses use Ink: https://github.com/vadimdemedes/ink
Re: Claude Opus 4.6
#225Re: Claude Opus 4.6
#226Earlier quoted context omitted.
Also explains why Claude Code is a React app outputting to a Terminal. (Seriously.)
There’s nothing wrong with that, except it lets ai skeptics feel superior
I thought this was a solid take
Re: Claude Opus 4.6
#227Earlier quoted context omitted.
There's no way they actually work on training this.
I suspect they're training on this. I asked Opus 4.6 for a pelican riding a recumbent bicycle and got this. https://i.imgur.com/UvlEBs8.png
Re: Claude Opus 4.6
#228Re: Claude Opus 4.6
#229The benchmarks are cool and all but 1M context on an Opus-class model is the real headline here imo. Has anyone actually pushed it to the limit yet? Long context has historically been one of those "works great in the demo" situations.
Re: Claude Opus 4.6
#230From the press release at least it sounds more expensive than Opus 4.5 (more tokens per request and fees for going over 200k context). It also seems misleading to have charts that compare to Sonnet 4.5 and not Opus 4.5 (Edit: It's because Opus 4.5 doesn't have a 1M context window). It's also interesting they list compaction as a capability of the model. I wonder if this means they have RL trained this compaction as o…
> From the press release at least it sounds more expensive than Opus 4.5 (more tokens per request and fees for going over 200k context). That's a feature. You could also not use the extra context, and the price would be the same.