Earlier quoted context omitted.
I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…
"It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years" What insight do you have to make this claim?
Claude Opus 4.8
401–410 of 1001 posts
Re: Claude Opus 4.8
#402Does anyone troll these releases and cherry pick random metrics other companies would cherry pick to show how amazing their models are? There's like 8 million benchmarks. Every release, every model randomly picks 5-10 where they win in everything except 1, to make it look like they aren't randomly cherry picking benchmarks they probably benchmaxxed for.
https://arena.ai/leaderboard - I’ve found this company is a pretty good ranker - not sure their exact methodology but during day to day programming with Claude / gpt models I’ve felt qualitatively what they report
Re: Claude Opus 4.8
#403I generated pelicans riding bicycles on both thinking level low and thinking level high: https://gist.github.com/simonw/68560eddb0b268a8417f80ceb7304... The high one is notably better - the bicycle frame is the correct shape, unlike thinking level low. For comparison, here's Opus 4.7: https://gist.github.com/simonw/afcb19addf3f38eb1996e1ebe749c...
I find the most miraculous thing about 4.7 to be that the pelican is facing left, wonder why the right facing everything is so ubiquitous in these images.
Re: Claude Opus 4.8
#404Re: Claude Opus 4.8
#405A rambling comment: I think this is the first time we've had a third minor version bump on a frontier Anthropic model. (I count the 0.5s as major here, because they've been issued non-sequentially and also corresponded to massive capability leaps, eg, Sonnet 3.5, Opus 4.5). So now the Opus 4.5 family has successors 4.6, 4.7, and 4.8, each posting fairly modest claimed gains. My own experience w/ 4.6 and 4.7 are that…
I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…
We have so many ways of optimizing:
- continusly creating more and better training data
- increasing parameters to 20/50/100TB
- We still wait for Mythos access
- We still wait for Mythos distilation (i haven't heard any rumors or so that there is a distilled version of Mythos out)
- Reinforcment learning and evolutionary algortihm only started to appear
- If a small 30GB Model can do stuff, these models can also be used as teachers for the big ones
- We have not seen yet specialized models at all. Like a coding java german expert model. Why? Even with MoE architecture, you still need to have these layers around
- Research for Diffusion and other models is still in progress
- Nvidia just announced/showed a 7x speedup on inferencing for Nemotron
- Multitoken prediction became available just a few weeks ago
- Compute gets only in a range were they can do a lot more and cheaper experiments (see Google IO 2026 announcement)
- World models are showing great progress and we do not know yet what they will bring to the table
- They are probably not finetuning/fixing all areas in parallel. I would argue that Anthropic focuses most of its efforts into coding and agentic. Google for sure does subagent and agentic optimizations too. Plenty of areas are just not touched i would say because they don't have the capacity
- We see more and more mulit modal models (these also consume compute)
- N-Gram paper and co i have not seen all of these things in chinese open models
- We don't even know yet what Meta is doing, but we do know they restarted their efforts again
- Anthropics models got a lot better benchmark wise for dening non sense asks. They do learn how to get rid or reduce hallucinations
- We are in the middle of the biggest Reinforcement loop whith all the training data we give them day to day and its not clear at all if they already use these models in thir training and at what stage.
- We do expect bigger models to be able to comprehend deeper concepts / broader code bases. Big companies with huge code bases probably are waiting for this
- Thre will be also continues progress in harnesses which in it alone is not part of the LLM progress (fair) but these harnesses do get better when you finetune a model to be optimized for a harness
- ChatGPTs Image model 2.0 got relevant better and came out just a month ago
I suspect, based on hardware requirements and progress on hardware infrastructure alone, that the industry wants to go to 100t models and we do not know yet what this will mean. I could see that we might skip normal transformer and find relevant other architectures.
Just a week ago there was a research paper about parallel input and output streams which has not been explored enough.
There was also a research paper were they showed that a LLM can compute things. This will take time to see were this leads to.
I don't think the focus on GRAM and facts is so relevant. Its about context and context handling not just some facts.
Re: Claude Opus 4.8
#406There is an obvious shift in sentiment amongst users, at least here in the US. I feel it myself, even as a proponent of AI tools, the bloviating and language that these companies use in these release articles are starting to wear thin on my patience. Its possible we might just be witnessing a shift in fashion, where this type of sentimentality was more acceptable when it was novel and new, but now it just appears out…
Re: Claude Opus 4.8
#407I don't know why the world is so happy about this when we should actually say stop.
Re: Claude Opus 4.8
#408Frontier models are mostly past the point of human ability to discern whether they are actually better or worse than predecessors and competitors. I suspect the benchmarks may also be saturated, or at least past their usefulness. I personally feel that Anthropic doesn't understand what this means for the frontier labs, and moreover that they might be the only frontier lab that doesn't. 1. Google dropped Gemini 3.5 Fl…
No idea why you’d say they have critically underinvested in product when Claude Code dominates and they’ve also released popular tools like Cowork and integrations for Microsoft products at an incredibly rapid pace.
Cost is becoming more of a factor, and no doubt they’ll work on that. There’s no reason to think they won’t be able to release cheaper models if they optimize for that rather than improving performance.
Re: Claude Opus 4.8
#409This is the first time I saw a model pop-up on HN and didn't really care. Model exhaustion? It looks interesting but not exciting. While I'd normally _love_ incremental improvements --- I think the recent ones are far too minor to get excited about or change up a workflow. Besides, benchmarks tend to exaggerate the gap between versions. At this point I'd almost rather Anthropic wait and really wow us with a 5.0 relea…
Re: Claude Opus 4.8
#410A rambling comment: I think this is the first time we've had a third minor version bump on a frontier Anthropic model. (I count the 0.5s as major here, because they've been issued non-sequentially and also corresponded to massive capability leaps, eg, Sonnet 3.5, Opus 4.5). So now the Opus 4.5 family has successors 4.6, 4.7, and 4.8, each posting fairly modest claimed gains. My own experience w/ 4.6 and 4.7 are that…
This felt particularly visible during the 4.6 when people said that 4.6 felt dumber and I remember someone doing some analysis and it sort of proved that models were getting dumber over time.
This has both benefits of costing less for the company to run while taking a standard subscription but also, at the same time, making the next model when it drops to public to "feel" more good comparatively.
Again, I am not sure if this is the case or not but merely proposing something that I feel like it might be in the possibility of realm.