The signal here is tokeneconomics are very real, price vs performance is starting to be a consideration even at the bleeding edge labs. maybe a subtle indication scaling is not all that is needed since if AGI was around the corner leading labs would still be incentivized to pour all resources into larger (smarter - or maybe not?) models
Claude Opus 5
261–270 of 1001 posts
Re: Claude Opus 5
#262Noticed none of the comparisons mention Kimi K3. Is there a comparison chart?
That's by design. Anthropic wants to make open-weight models illegal (not my speculation -- Dario explicitly said so), so I assume they don't want to give them any undue attention.
Re: Claude Opus 5
#263What's the point of 150 pages description of a model that's going to be replaced in a couple months? Who even reads this? I know it's cheap to generate text with LLMs, but this is just noise at this point.
I actually do read them. Not in severe detail, but not casually either. 150 pages is really not very long and there doesn't seem to be too much bloat. (I would cut out the moral personhood stuff but that's a political/ideological thing). This is snarky but I am grumpy: I wonder if there's a correlation between me refusing to use LLMs and me being happy to read a novella-sized PDF about them.
Maybe it's just me, but 150 pages is like third of a good book. Quite long. And it's full of LLM slop, they did not even bother to remove the em dashes.
Re: Claude Opus 5
#264Re: Claude Opus 5
#265Looking at all these releases it’s not a surprise that model routing is the fastest growing segment in AI right now. There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/outp…
Who are the customers though? Honest question, I'd like to understand it. For me, anything other than current best available SOTA for any task is unacceptable. The only routing rule I need is "the most powerful model I still have flat-priced quota available for". I mean, why settle for less?
Re: Claude Opus 5
#266Earlier quoted context omitted.
Because their model previously got blocked by the government for this and they don't want a repeat?
5.6-Sol is a lot more permissive than Opus/Fable even w/ CVP (once you sign your soul away to Palantir via Persona, anyway), while maintaining better capabilities
Re: Claude Opus 5
#267Earlier quoted context omitted.
I suspect they make a big model first. In this case it's Fable. Then they run the shrinker steps to make Sonnet and Opus. Sonnet is smaller, takes less time to make, so it got released first. Opus needed few more weeks to cook. With this iteration they had a delay because when the Mythos was ready they had some sort of "Oh shit" moment and spent half a year adding safety guards to it. Then slowly rolled it out, but g…
But I wonder how they were able to release Sonnet 5 during the period when even people inside Anthropic were legally barred from using Mythos/Fable?
So I'm assuming at least a subset of employees could continue using the models during that time.
Re: Claude Opus 5
#268Earlier quoted context omitted.
So the rumors were right, Opus 5 was indeed being polished up for release. Huge improvements in GDPval-AA v2 too -- great for some of the knowledge work-based agentic workloads I run. Also glad they still kepy Fable 5 on "credits only" access. I think we're going to start seeing model providers gate top-of-the-line models behind pay-as-you-go API rates/credits while subsidizing other models on monthly subscriptions.
Fable 5 is still included in Max subscriptions!
Re: Claude Opus 5
#269Anyone has an insight into how much money labs are putting into benchmarks? Just Arg-AGI-3 is quoted above 20K USD and footnote says average of 5 runs (!!). Likely just a drop in the bucket to the training budget but still..
Re: Claude Opus 5
#270Earlier quoted context omitted.
Go read the safeguards section in the report and you will realize why that is. These models are heavily as safeguarded and that was the initial reason why they said they couldn't and haven't released Mythos because that model is the one without the safeguards. OpenAI is did the same thing when they announced a model without safeguards broken into HuggingFace servers.
Yes, this makes a lot of sense, but it’s just very amusing to see. 2 months ago, the world was about to end, now not so much.
i think we'll see one of the fastest deflations in history post anthropic/oai ipo