Live data from Hacker News

Claude Opus 4.8

anthropic.com

951–960 of 1001 posts

Re: Claude Opus 4.8

#951

Earlier quoted context omitted.

I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…

Took me a while to find what you were referring to by gram. Arxiv paper from 9 days ago that's not properly indexed by search engines. (G)enerative (R)ecursive re(A)soning (M)odels. They really wanted the acronym. https://arxiv.org/html/2605.19376v1

Let's not forget about Yann LeCun's current area of research that's completely different from LLMs: Joint Embedding Predictive Architecture (JEPA)

If he gets that style to be more efficient (they're already competitive) it'll completely kill off LLMs

https://openreview.net/pdf?id=BZ5a1r-kVsf

Re: Claude Opus 4.8

#952
post #77

A rambling comment: I think this is the first time we've had a third minor version bump on a frontier Anthropic model. (I count the 0.5s as major here, because they've been issued non-sequentially and also corresponded to massive capability leaps, eg, Sonnet 3.5, Opus 4.5). So now the Opus 4.5 family has successors 4.6, 4.7, and 4.8, each posting fairly modest claimed gains. My own experience w/ 4.6 and 4.7 are that…

I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…

>As far as reasoning is concerned, with the recent GRAM release

Graphic RAM?

Re: Claude Opus 4.8

#953

Earlier quoted context omitted.

I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…

surely training also gets cheaper so justifying it becomes easier? i think it'll be more like we get 1-10T models and then distill those down into smaller models, though It seems like the best small models today are all distilled from bigger models Moreover, I hypothesize Claude Opus 4.7 and now 4.8 are a distillation of Claude Mythos

That's the impression I got too, it seems closer to what the marketing has told us about Mythos than 4.6/4.7 were.

Re: Claude Opus 4.8

#954

Earlier quoted context omitted.

I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…

> Google, OpenAI, Anthropic could train a 30B GRAM-based model in days - and it could potentially have better local reasoning than the best model available today at >1T param I agree but with their urgent IPO-driven need to keep increasing prices, the frontier vendors now have every incentive maintain the perception that frontier performance requires endless >$200K racks of unobtanium GPUs and RAM. While they'd love…

> While they'd love to reduce their actual costs, they'd only want to do it to the extent they are certain they can keep it secret.

So you are saying that frontier AI labs are spending billions of dollars on datacenters as a form of marketing. And they are colluding to hide the fact that they don't need to.

Of course they profit more if they are in front, but bleeding money to pretend to be in front is not a winning strategy. They can't fool the market if their models are not actually better, and they know this.

Re: Claude Opus 4.8

#955
post #909

Earlier quoted context omitted.

pretty spot on. In my experience, Opus 4.0 was fantastic, major jump from 3.7. it was creative, super slow and expensive, and would sometime forget what it was doing, but it was getting the job done. 4.1 they made it much faster, so a lot of infra improvements. 4.5 was the time it could work on longer task, didn't make a lot of obvious mistakes of 4.0, and i think this was about the time the opus went mainstream, and…

> "4.6 was such a bad model," It's just amusing reading all these posts with different viewpoints, just in this thread there are multiple people saying 4.6 was so much better than 4.7 and that they switched back to 4.6.

I also find it amusing. I also heard a lot of "4.7 is garbage, everybody hates it". Shows you how important proper validation techniques are, not just gut feeling.

Re: Claude Opus 4.8

#957
post #942

As if choosing a model to use on its own is not hard, offering six levels of "effort" (quite a vague term as well), low, medium, high, xhigh, max, ultracode (?!?!) is really making comparisons next to impossible when people using the same model can have vastly different experiences. What exactly is the diff between high and xhigh? Or xhigh and max? This is definitely too granular and it seems Anthropic took OpenAI's…

Feels like one of those knobs they throw in to make performance feel like a user skill issue instead of a product issue.

Why doesn't it know how much effort to use? How do I know how much effort to use? It's a mystery.

Re: Claude Opus 4.8

#958

I find it freaky how you notice the language change between models. Some words which pop up now all the time, that I don't remember reacting to with previous models, such as "honest(ly)" and "load-bearing". Feels like a new AI smell, like em-dashes or "it's not just x, it's y".

I’ve never asked Claude to re explain things so many times with this. The language it chooses is bizarre, not quite technical enough to be precise, and too hand wavy to be useful.

There also seems to be other issues with the sub agents and not inheriting memory.

E.g. I’m working on an analysis, absence of row implies a zero (it’s sparse).

Every time it checks the results it presents this missing data bug, but it’s not a bug, and I’ve explained it 10 times.

I mean, at finding issues it’s been great. I can ask it to look in other repos and its finding cross repo behaviour or subtleties much faster.

But it’s wasting more tokens reporting caveats I’ve already explained. And it’s suggesting stopping analysis when there’s still work to be done.

So overall I’d say it’s mixed. Feels better at code but worse at my explicit preferences and objective analysis.

Re: Claude Opus 4.8

#959
post #77

A rambling comment: I think this is the first time we've had a third minor version bump on a frontier Anthropic model. (I count the 0.5s as major here, because they've been issued non-sequentially and also corresponded to massive capability leaps, eg, Sonnet 3.5, Opus 4.5). So now the Opus 4.5 family has successors 4.6, 4.7, and 4.8, each posting fairly modest claimed gains. My own experience w/ 4.6 and 4.7 are that…

I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…

For that matter, we may have models/tooling that are smaller that are designed for say identification model first, then handoff to a context specific model that is optimized for a specific domain... where the two calls through tooling are more optimal than a single call to a much large model. We're already kind of close to this with how the likes of claude code work with handoffs to other tools/modules.

I can see a LOT of room to explore and partition domains into more specified models still.

Re: Claude Opus 4.8

#960

Earlier quoted context omitted.

I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…

Small models don't have enough parameters to memorize the entire internet . For very common prompts you don't notice that, but when you rely on some niche knowledge that might only appear once in the entire web, a single blogpost, a single github issue, a single pdf, you need to be lucky enough that the agent runs a web search AND it returns what you need. Even as humans there's so much knowledge out there that exist…

Exactly, as humans you won't know everything... but you CAN know enough to roughly classify what to "google" for... And if you can google a problem summary, you could identify from a select list of domain specific AI models to use one or more to aggregate work results. And if a person can do that, a model can be trained to do/leverage the same.

You can have a domain limited classification model that then passes the query/work to best match model(s) that do the work... then rollup the results. Basically two very cheap requests instead of one much more expensive one.

Post reply on HN