Live data from Hacker News

Claude Opus 4.8

anthropic.com

671–680 of 1001 posts

Re: Claude Opus 4.8

#672
post #73

I generated pelicans riding bicycles on both thinking level low and thinking level high: https://gist.github.com/simonw/68560eddb0b268a8417f80ceb7304... The high one is notably better - the bicycle frame is the correct shape, unlike thinking level low. For comparison, here's Opus 4.7: https://gist.github.com/simonw/afcb19addf3f38eb1996e1ebe749c...

Glad to see that the "high thinking" level adds a helmet. Always a smart choice.

And yet some people doubt Anthropic's commitment to AI safety

Re: Claude Opus 4.8

#673
post #574
post #485

Earlier quoted context omitted.

What is ultracode mode?

It's a combination of reasoning effort (max) + enabling workflow that orchestrates multiple sub-agents. After some interrogation, here's how it organized the work: 1. Design workflow (rts-game-design, 11 agents, ~13 min) ran first, produced SPEC.md + DESIGN.md: 1.1. Proposals (3 parallel agents): each designed a complete RTS from a different philosophy 1.2 Judge (1 agent): evaluated all three and synthesized one unif…

seems like a rube-goldberg esque way to consume 10x tokens. is this really where the industry is heading?

Re: Claude Opus 4.8

#674
post #394

Frontier models are mostly past the point of human ability to discern whether they are actually better or worse than predecessors and competitors. I suspect the benchmarks may also be saturated, or at least past their usefulness. I personally feel that Anthropic doesn't understand what this means for the frontier labs, and moreover that they might be the only frontier lab that doesn't. 1. Google dropped Gemini 3.5 Fl…

This post is proof that people will complain about anything, even if its the most successful startup of the past decade.

You're not successful until you exit. And, of course, there's always room to be more successful.

Re: Claude Opus 4.8

#675

"Users will find Opus 4.8 to be a modest but tangible improvement on its predecessor." This is a refreshing attitude! I've also verified that you can now turn off adaptive thinking in the web UI, which is great. I've had a lot of problems with thinking not triggering and the model producing sub-par output. Glad we can finally turn it off. (I hope being able to turn off adaptive thinking is new, if I could have turned…

I was hoping that the web UI would be better -- I like Anthropic better than OpenAI from a values perspective and want to use their products, but ChatGPT in thinking mode has been just vastly better than claude.ai.So my fingers were crossed that these changes would bring it up to par. But trying it out... alas, no. Simple factual questions where ChatGPT would go do a quick search and get the facts and report them bac…

What are some examples?

Re: Claude Opus 4.8

#676

Given DeepSWE just blew apart the SWE-Bench Pro benchmark and handed a 14-point lead to GPT-5.5, it looks pretty bad that they've listed SWE-Bench first in the model release and no DeepSWE. Like, this isn't obviously an answer. Or maybe it is, but publish the DeepSWE numbers so we can see for ourselves.

This is a terrible benchmark. It literally tests the models on their ability to track shifting line numbers. If they cannot keep up, no amount of abstract reasoning can redeem them.

Re: Claude Opus 4.8

#677

All I need for Christmas is a Claude that doesn't spit out so many em dashes.

And that doesn't use "worth flagging" and "load-bearing" in every other sentence.

You're absolutely right - and I should have tempered that behavior. When the next version lands you get much better responses. Not just trite analogies. Really well spoken responses that earn their keep.

Re: Claude Opus 4.8

#678
I haven't had the best experience with 4.7 and it felt like a substantial debuff. I've even ended up moving a lot of review to codex just because 4.7 was so dense.. Here's to hoping they figured it out since I'm not entirely sure but I would have to guess that they were experimenting with making the model lighter (although I have no concrete evidence of this).

Re: Claude Opus 4.8

#679
next (or maybe current) frontier of competition may not be the model, rather the harness and how much unique advantage a lab-created harness can beat 3rd-party harness.

Re: Claude Opus 4.8

#680

Earlier quoted context omitted.

I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…

Took me a while to find what you were referring to by gram. Arxiv paper from 9 days ago that's not properly indexed by search engines. (G)enerative (R)ecursive re(A)soning (M)odels. They really wanted the acronym. https://arxiv.org/html/2605.19376v1

It is the 3rd list on Kagi when searching "gram models"
Post reply on HN