There is an obvious shift in sentiment amongst users, at least here in the US. I feel it myself, even as a proponent of AI tools, the bloviating and language that these companies use in these release articles are starting to wear thin on my patience. Its possible we might just be witnessing a shift in fashion, where this type of sentimentality was more acceptable when it was novel and new, but now it just appears out…
Watch Christopher Olah bloviate at the Vatican during the Magnifica Humanatis launch. It's truly nauseating. I've never seen such a ridiculous speech in my life. Between him and the CEO, I'm starting to understand the level of arrogance these people are capable of.
Claude Opus 4.8
791–800 of 1001 posts
Re: Claude Opus 4.8
#792Re: Claude Opus 4.8
#793My fav coding benchmark for frontier models is to build a simple RTS game in one file (js/html/css). Claude Code with Opus 4.8 in ultracode mode nailed it, the best result so far: https://bsky.app/profile/senko.net/post/3mmwnrkwboc2v The prompt was: Create a simple but functional real time strategy (RTS) game similar to old WarCraft, StarCraft or Command & Conquer games. The player should be able to build buildings,…
Re: Claude Opus 4.8
#794Still not worth the cost over GPT 5.5. Anthropic better start improving their speed+costs, or they're going to lose an incredible amount of business. And no, fast mode is not something any sane person will ever use. 6x the cost for 2.5x the speed, what a joke...
Re: Claude Opus 4.8
#795Re: Claude Opus 4.8
#796> how many days in the week have the letter d in them?
> Two - Sunday and... wait, let me actually check. Monday, Tuesday, Wednesday, Thursday, Friday, Saturday, Sunday. The ones with a "d": Wednesday, Thursday, and Sunday all have one. Monday too. So that's Monday, Wednesday, Thursday, Sunday - four days.
Re: Claude Opus 4.8
#797Rollout has been a little suspect. Hope it gets better.
And after that asked some questions that it already had answers to.
Started a brand new session and it's been OK since. Only drawn one silly conclusion so far, which I nudged it away from.
Re: Claude Opus 4.8
#798Earlier quoted context omitted.
I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…
And anyway, with quantum, there will be no need for frontier companies as you might be able to even run a 1T param model on a consumer quantum computer.
Re: Claude Opus 4.8
#799Re: Claude Opus 4.8
#800Still feels like even with Max mode it doesn't think reasonably long, at least ChatGPT Pro thinks longer.