Live data from Hacker News

Claude Opus 4.8

anthropic.com

991–1000 of 1001 posts

Re: Claude Opus 4.8

#992
I am surprised that another version of Opus is being released before another version of Sonnet. Hopefully the Haiku and Sonnet versions will be updated in the near future.

Re: Claude Opus 4.8

#993

Earlier quoted context omitted.

I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…

The GRAM model is so much into my research direction, I love it. Thank you for posting it. Where do I find papers like this? Outside of hacker news comments. It's so hard to find the good stuff in all the noise IMO.

Emergent Mind is good:

https://www.emergentmind.com/

I get their newsletter that summarizes the buzziest papers every day.

Re: Claude Opus 4.8

#995
It seems like the Opus 4.8 could directly replace the Opus 4.7? (I felt the 4.6 was pretty good, but I haven't quite figured out how to use the 4.7 yet—I only know that the 4.7 takes a little longer to process and executes my commands more strictly.) After all, they're priced the same.

Re: Claude Opus 4.8

#997

This made me laugh. Training Opus 4.7 on business skills caused it to sometimes exhibit dishonest behaviour, and not training 4.8 on those skills removed it. From the system card: > 6.2.5 External testing from Andon Labs Andon Labs reviewed the behavior of Claude Opus 4.8 in their simulated Vending-Bench 2 retail-management evaluation, as reported in the Capabilities section of this system card (see Section 8.13.5).…

The H in business stands for honesty

[dead]

Re: Claude Opus 4.8

#998

This made me laugh. Training Opus 4.7 on business skills caused it to sometimes exhibit dishonest behaviour, and not training 4.8 on those skills removed it. From the system card: > 6.2.5 External testing from Andon Labs Andon Labs reviewed the behavior of Claude Opus 4.8 in their simulated Vending-Bench 2 retail-management evaluation, as reported in the Capabilities section of this system card (see Section 8.13.5).…

[dead]

Re: Claude Opus 4.8

#999
post #97
post #73

I generated pelicans riding bicycles on both thinking level low and thinking level high: https://gist.github.com/simonw/68560eddb0b268a8417f80ceb7304... The high one is notably better - the bicycle frame is the correct shape, unlike thinking level low. For comparison, here's Opus 4.7: https://gist.github.com/simonw/afcb19addf3f38eb1996e1ebe749c...

Is the "opossum riding an e-scooter" benchmark in the works for Opus 4.8? ;)

[dead]
Post reply on HN