This is the first time I saw a model pop-up on HN and didn't really care. Model exhaustion? It looks interesting but not exciting. While I'd normally _love_ incremental improvements --- I think the recent ones are far too minor to get excited about or change up a workflow. Besides, benchmarks tend to exaggerate the gap between versions. At this point I'd almost rather Anthropic wait and really wow us with a 5.0 relea…
Claude Opus 4.8
911–920 of 1001 posts
Re: Claude Opus 4.8
#912Earlier quoted context omitted.
I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…
It is fascinating to me to see a new product category that improves so vastly year-after-year, where people commonly state that this is now the peak already. I couldn’t even imagine having to go back to a model from 12 months ago, much less 24 months ago. GPT-5.5 is so much better than GPT-4o that it sure seems like they keep finding new juice to squeeze. This is like going from dialup internet to DSL and acting like…
The difference in progress in smaller models is far more impressive.
Compare Gemini 3.5 Flash to a ~16B parameter model from 24 months ago.
Compare GPT-5.5 to a frontier model 24 months ago.
Yes, GPT-5.5 got better. At orders of magnitude smaller parameter sizes (when factoring in ACTIVE parameters) the increase is far more pronounced.
Re: Claude Opus 4.8
#913Re: Claude Opus 4.8
#914Earlier quoted context omitted.
I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…
The GRAM model is so much into my research direction, I love it. Thank you for posting it. Where do I find papers like this? Outside of hacker news comments. It's so hard to find the good stuff in all the noise IMO.
I got it from my Google News recs on my phone, because I've been watching a bunch of videos on YouTube about LeCun's ideas on World Models and JEPA (I think).
Re: Claude Opus 4.8
#915Does anyone troll these releases and cherry pick random metrics other companies would cherry pick to show how amazing their models are? There's like 8 million benchmarks. Every release, every model randomly picks 5-10 where they win in everything except 1, to make it look like they aren't randomly cherry picking benchmarks they probably benchmaxxed for.
It's interesting they only included 6 metrics this time. Opus 4.7 had 12, and 4.6 had 13. Of the metircs they reported for 4.7, for 4.8 they excluded BrowseComp, CharXiv Reasoning, CyberGym, GPQA Diamond, MCP Atlas, MMMLU, SWE-bench Verified. The last 4 were almost always mentioned in previous Opus releases.
Re: Claude Opus 4.8
#916Meanwhile Deepseek is cutting inference costs to mere cents. Thats the real AI revolution for you.
It is basically indistinguishable from sonnet. At this point my own prompts, AGENTS.md, background docs and so on matter a great deal more than the differences between models.
And deepseek v4 flash (the sonnet comparable) costs 3% of what sonnet does.
Re: Claude Opus 4.8
#917Meanwhile Deepseek is cutting inference costs to mere cents. Thats the real AI revolution for you.
Yes I switched from claude code to opencode with deepseek recently. It is basically indistinguishable from sonnet. At this point my own prompts, AGENTS.md, background docs and so on matter a great deal more than the differences between models. And deepseek v4 flash (the sonnet comparable) costs 3% of what sonnet does.
Re: Claude Opus 4.8
#918> Not only that, but we plan to release a new class of model with even higher intelligence than Opus. As part of Project Glasswing, a small number of organizations are currently using Claude Mythos Preview for cybersecurity work. Models of this capability level require stronger cyber safeguards before they can be generally released. We’re making swift progress on developing these safeguards and expect to be able to b…
Re: Claude Opus 4.8
#919"Users will find Opus 4.8 to be a modest but tangible improvement on its predecessor." This is a refreshing attitude! I've also verified that you can now turn off adaptive thinking in the web UI, which is great. I've had a lot of problems with thinking not triggering and the model producing sub-par output. Glad we can finally turn it off. (I hope being able to turn off adaptive thinking is new, if I could have turned…
Yes, modest but tangible improvement - same modesty does not apply to the cost: https://artificialanalysis.ai/models/capabilities/coding#cod...
Re: Claude Opus 4.8
#920Earlier quoted context omitted.
I think 4.7 was an awful model in actual use. I never got anything out of it and it was frustratingly weird. This feels more like an attempt to course correct and isn't a real bump
I think they overtrained on scientific papers or such as it would spout really sophisticated sounding nonsense with a ton of complicated verbs and adjectives. 4.6 was definitely better in that regard. The more I use these tools the more I think they’re not actually that revolutionary. I mean it’s still amazing what they can do but they have very clear limitations it seems.
Frustrating because if I have a tool, I expect a tool to do what I tell it to do. Tools shouldn't have any opinions on how they should be used