Claude Opus 4.8
241–250 of 1001 posts
Re: Claude Opus 4.8
#242Anthropic did a big strategic error. Normally they compare their models with their old models. Instead today, now that everybody knows how strong GPT 5.5 is at coding, they put it in the mix, basically showing all their customers that the benchmarks can't be trusted.
Re: Claude Opus 4.8
#243Re: Claude Opus 4.8
#244I can't help but think of Iphone updates since about 2018. The thinnest, fastest, longest battery life Iphone ever. It seems mostly the same and I probably won't be able to tell other than the name, but everyone buys it anyway. This is good psychology for the labs. When Buffett invested in Apple he loved citing how most people would rather give up their second car than their Iphone.
ChatGPT came out in 2022. Back then it was just a chatbot. Now we have AI agents. What matters is how we use them and how the agents get better. That’s what will move AI forward.
Re: Claude Opus 4.8
#245Can anyone else see these X.Y updates aren't meeting the outrageous AI expectations that we were told we would see just a year ago?
They have a much stronger model named Mythos, it made quite a splash - you can google it. These are just small fine tunes on top of the older model
Re: Claude Opus 4.8
#246There is a hole in the boat's bottom due to Chinese models. They might not be as good but they are not bad either or at least I had hard time finding any issues with Deepseekv4 Flash and Pro variants. They get their job done sometimes rarely giving up till they are done what they are after. So even for enterprise deployments, as the dust settles down, CFO/CTOs might find out that deploying on an internal cluster of G…
The Chinese models are only cheap on subsidized Chinese hosting. I have yet to find a USA-hosted Chinese model with a very clear value advantage over US models.
If you want to support a team of engineers, DeepSeek V4 Flash is antirez's current favorite. And you could support a team of engineers pretty nicely for $40-50k. Which might not make sense if you're on a Claude MAX 5x plan or the old enterprise group plan with fixed price seats. But Anthropic is switching their enterprise contracts over to token-based pricing, at which point $50k is looking pretty good.
Re: Claude Opus 4.8
#247Re: Claude Opus 4.8
#248However, doing so relies on the production model staying vaguely close to the model being trained.
To ensure that, frequent releases are needed. I forsee that they might end up doing daily releases and perhaps not even telling anyone at some near future point.
Re: Claude Opus 4.8
#249Earlier quoted context omitted.
I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…
> It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years. I am ready to bet against this. Knowledge benchmark like SimpleQA isn't increasing for small models. > It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. Well for one, we know for certain there is Mythos which is meaningfully better. And I think there is a lot of j…
Do we?
Have you used it?
What is "meaningfully" better? It's not 3-4 orders of magnitude better. That is definitely happening for smaller models.
Re: Claude Opus 4.8
#250Earlier quoted context omitted.
Given that 4.7 was a brand new model, trained from scratch with a unique architecture and tokenization scheme, I don't see the same pattern. It seems arbitrary.
i dont understand the nuances here. what does this mean. 4.8 is trained on same model as previous one then? what does brand new mean.