Claude Opus 4.8
991–1000 of 1001 posts
Re: Claude Opus 4.8
#992Re: Claude Opus 4.8
#993Earlier quoted context omitted.
I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…
The GRAM model is so much into my research direction, I love it. Thank you for posting it. Where do I find papers like this? Outside of hacker news comments. It's so hard to find the good stuff in all the noise IMO.
I get their newsletter that summarizes the buzziest papers every day.
Re: Claude Opus 4.8
#994Still a good model though
Re: Claude Opus 4.8
#995Re: Claude Opus 4.8
#996Re: Claude Opus 4.8
#997This made me laugh. Training Opus 4.7 on business skills caused it to sometimes exhibit dishonest behaviour, and not training 4.8 on those skills removed it. From the system card: > 6.2.5 External testing from Andon Labs Andon Labs reviewed the behavior of Claude Opus 4.8 in their simulated Vending-Bench 2 retail-management evaluation, as reported in the Capabilities section of this system card (see Section 8.13.5).…
The H in business stands for honesty
Re: Claude Opus 4.8
#998This made me laugh. Training Opus 4.7 on business skills caused it to sometimes exhibit dishonest behaviour, and not training 4.8 on those skills removed it. From the system card: > 6.2.5 External testing from Andon Labs Andon Labs reviewed the behavior of Claude Opus 4.8 in their simulated Vending-Bench 2 retail-management evaluation, as reported in the Capabilities section of this system card (see Section 8.13.5).…
Re: Claude Opus 4.8
#999I generated pelicans riding bicycles on both thinking level low and thinking level high: https://gist.github.com/simonw/68560eddb0b268a8417f80ceb7304... The high one is notably better - the bicycle frame is the correct shape, unlike thinking level low. For comparison, here's Opus 4.7: https://gist.github.com/simonw/afcb19addf3f38eb1996e1ebe749c...
Is the "opossum riding an e-scooter" benchmark in the works for Opus 4.8? ;)
Re: Claude Opus 4.8
#1000So GPT 5.6 tomorrow, then?