Live data from Hacker News

Why does Opus 5 feel worse to work with?

mun-logadan.github.io

821–830 of 915 posts

Re: Why does Opus 5 feel worse to work with?

#821

Earlier quoted context omitted.

>it feels like reading an impression of a literature book by a high school English class’s most overconfident student who’s only ever read LinkedIn-speak. Claude is very much the “stupid person’s idea of an intelligent person”[0] which, I suspect, is why it is so popular. It certainly explains why half the internet is huge chunks of Claude-authored gibberish copied and pasted and published. If people didn’t think it…

there's a wide array of assessments when it comes to reading comprehension. the one you refer to, the GRA, sets the 'sixth grade level' as whether or not a reader understands the author's main points, is able to answer conceptual questions related to the text, and then apply those to relevant situations. beyond this level is the ability to essentially be skeptical of a text and to know how to critically analyze it. s…

> I'll also say that I think Claude sounds the way that it does because it, like many other LLMs, are RLHF trained largely by lowly paid gig-workers, many of them ESL speakers. if their trainers were, for example, dedicated and highly trained academics, scientists, and other researchers, you'd likely see a lot more concise and more importantly skeptical reasoning and responses. but that won't happen in our current reality of capitalist-driven development so we get encoded solutions like MoE that still largely depend on the messy, imprecise RLHF training at baseline

No really, that's not particularly accurate, they use so much gig work because no-one else wants to work for them not because they would be unwilling to pay a little extra, or only want the absolute cheapest labor they can get on the planet.

They want senior white collar professionals and scientists and researchers especially since these companies already on some level believe their models are as good as any senior employee in any field (it's probably the models generating text saying that, but that's besides the point). But who's going to work on contract for a company that wants to automate them out of a job? Realistically no-one unless they get some shares in the thing that will destroy their future earnings potential and ability to control their own destiny if it works out.

But they can find enough educated white collar professionals on unemployment or in unstable academic employment that will take an extra job on even if it's only 50 $/h or 70 $/h and compromise on any solitary they might have but the work output you get from that is only going to be as good as what you ask for, if they had better respect for the professions they want to automate, it would be better.

Like is that an acceptable wage in the US for difficult skilled work, not particularly but it's not rock bottom exactly, and it's not bad for other English speaking countries, working conditions and stated mission are more of an issue than being cheap.

Training pipeline on a modern LLM is also going to be quite indirect during the long tail of post training, and heavy on automated RL, the human feedback might end up getting used in the form of automated grading guidelines like what you did for research, with the same issues as that, compounded by the input being LLM generated and models being biased towards model output by default. It's more of a feedback on the loop rather than in the loop.

Re: Why does Opus 5 feel worse to work with?

#822
post #372

Earlier quoted context omitted.

CLAUDE.md is mostly powerless against the reinforcement learned crap. I'm up to three separate instructions telling it to cut out the hyper verbose, retelling history comments and it still writes them every time.

Yes. CLAUDE.md only works half the time, except in longer conversations, when it works about 10% of the time. Hooks are also useless in the sama manner, the agent learns to dodge “no comments” hooks (why is it adding them anyway?). Hooks to append text to your prompt reminding the agent of certain rules are useless. Claude does whatever it wants, when it wants, the way it wants

You could probably keep the Claude slop hidden and have a fresh model generate a paraphrased response for anything human visible and keep the Claude responses as thinking.

Re: Why does Opus 5 feel worse to work with?

#823
post #476

Earlier quoted context omitted.

CLAUDE.md is mostly powerless against the reinforcement learned crap. I'm up to three separate instructions telling it to cut out the hyper verbose, retelling history comments and it still writes them every time.

You would hope? Really really hope? that they could observe this, and target it? Like, Claude going off the rails isn't something that takes a lot of effort to demonstrate. Literally anybody with a CLAUDE.md has seen the behavior over and over and over. Hey Ants, can you maybe just not release the next version, no matter how good it seems on benchmarks, if it can't follow the goddamn instructions? Please? This seems…

Working as intended, the purpose of a system is what it does.

Re: Why does Opus 5 feel worse to work with?

#824
post #424

Earlier quoted context omitted.

I'm switching to GPT because of this. The prose is so much more legible. The only reason I keep using Claude Code is because the harness is the best IMO.

I was the same until I ran out of Anthropic tokens one day and used "Grok Build" which is their Claude Code clone. You can use config to point it any LLM API so don't need to use Grok, and I like the UI better too.

I'm not sure I could really live down using something branded with Grok but it makes sense Elon Musk at least shipped user facing products in the past, not surprised his company delivered something more usable.

Re: Why does Opus 5 feel worse to work with?

#825
post #428

Earlier quoted context omitted.

Everything that claude writes fits into the same aesthetic structure. The aesthetic is that of an expert slowly revealing an insight to the user. The actual content doesn't matter. - "Introduction that rephrases your prompt." - "3 paragraphs, with one section of bullet points" - "The Twist" - "The Bottom Line" It's really obvious once you see it. Every single prompt, from a quantum physics question to a mundane obser…

Yes! Ive started getting a feel for AI writing on blogs. It feels slightly verbose and involves "reveals" "It's not the naked man on your lawn waving a chainsaw that's scaring you. It's the burrito you ate for lunch: it went down easy, but now it's coming for you"

Probably trained on lots of clickbait.

Re: Why does Opus 5 feel worse to work with?

#826
post #294
post #175

Earlier quoted context omitted.

It seems to have a preference for speaking in poetic or highly expressively language, rather than precise and concise as most engineers like to talk. The amount of times I have to ask "precisely what do you mean by x?". It's kinda like that engineer that likes to throw around unnecessary technical jargon just to sound more inteligent, worse because at least you could kinda understand what the technical jargon dude wa…

Claude writes like a guy at a firm I used to work with in the 90s; he was my employer's "visionary"; he'd worked at a whole lot of different companies on both sides of the Atlantic in inexplicably high-placed roles given that he was often bluffing, and was considered a lucky hire of a rising star. He'd be called into meetings with high end clients to spout off. He really needed you to know he understood, but very oft…

Probably to a degree, I have found Gemini to be the least dis-likable of the models from the big 3 on that front. I wonder if the poor English comprehension of Deepseek-v4-pro and K3 is because of alleged distillation of Claude (speaking of why doesn't anyone distill openAI, are they just dramatically more competent at stopping API use that breaks their terms?).

V4-pro in particular seems very capable, but will just dramatically completely misunderstand user intent, it seems almost like it wasn't trained at all on non LLM generated instructions mid conversation.

Re: Why does Opus 5 feel worse to work with?

#827
post #175

Earlier quoted context omitted.

It seems to have a preference for speaking in poetic or highly expressively language, rather than precise and concise as most engineers like to talk. The amount of times I have to ask "precisely what do you mean by x?". It's kinda like that engineer that likes to throw around unnecessary technical jargon just to sound more inteligent, worse because at least you could kinda understand what the technical jargon dude wa…

I asked some AI-using compatriots a while back who were complaining about this, 'isn't it doubling down on bullshitting you?' and got some pushback along the lines of 'it isn't a person therefore doesn't have dark motives like that therefore can't be doing that to us'. Didn't convince me. I think bullshitting like this can be a behavior, not just the intention of a human. If it's blowing a lot of smoke to use fancy w…

They have written like that when the models were much less capable, my hypothesis is this is an example of model collapse happening ever since LLM training leaned in heavily into RL and a result of training on model output the developers are uninterested in correcting since they want ASI not a somewhat useful AI coding tool that supplements humans without replacing them in the economic system.

Re: Why does Opus 5 feel worse to work with?

#828
post #674

I'm with the author and others in this comment thread, speculating that effectively the balance has tipped to where humans are no longer the target audience of post training - other agents are. Whether it's through the reasoning / CoT, or whether it's in handing off to subagents etc, the focus has moved to agents communicating in "agent-speak" to themselves or other agents. And human niceties are just kind of, noise…

Every round of models (plus all the secret tweaks) require new strategies to stay afloat as a human. My new tactic for Fable and Opus is to give them a line limit, both during planning and code creation. It os amazing how well that works for keeping them on task and avoiding premature optimization, pointless tests or any of those "robustness" ideas that are not planned or asked for.

Before I even look at a PR of sol I ask it to justify the loc. Quite often it comes back suggesting things it could simplify

Re: Why does Opus 5 feel worse to work with?

#829

Earlier quoted context omitted.

the phraseology is unbearable, it speaks like some kind of pretentious dude from a software engineering discord or something, littered with lingo and catch phrases I try to push through but it's insufferable

It speaks like a Senior Staff Software Engineer who was somehow hired into that title with 6 months of work experience.

I suppose it actually is exactly that, bar not being a human being that experiences anything

Re: Why does Opus 5 feel worse to work with?

#830
post #12

The single biggest annoyance with Opus 5 is that it writes too elliptically. Sentences that orbit a point, then jump to it like it's a revealed insight. Unnecessarily abstract phraseology. Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice, especially when it helps construct a sentence where the real action can 'land' like a surprise at the end. It is definitely mo…

Why are you so distracted by style? I see it too, it could be improved, it might be improved, but I don’t care as long as it get shit done. And it does, tons of shit gets done. And the communication is usually more informative than how coworkers document their work. Can it be improved? Yes. Does that mean I can’t use it? No.
Post reply on HN