Live data from Hacker News

Why does Opus 5 feel worse to work with?

mun-logadan.github.io

751–760 of 915 posts

Re: Why does Opus 5 feel worse to work with?

#751

A gem Opus 5 gifted to me today: "A devastating pair of findings, and the first is beautiful in a way worth naming: the anti-vacuity floor is what blinds the gate to a vacuous case."

You're right that I shouldn't have used the forbidden words; that's entirely on me. What I also did was identify the plausible-gate issues. Want me to tackle them next?

Re: Why does Opus 5 feel worse to work with?

#752
post #12

The single biggest annoyance with Opus 5 is that it writes too elliptically. Sentences that orbit a point, then jump to it like it's a revealed insight. Unnecessarily abstract phraseology. Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice, especially when it helps construct a sentence where the real action can 'land' like a surprise at the end. It is definitely mo…

Replying here without having read any of the sub-replies so I apologize if this is a repeated theme.

I have explicit markdown about telling the model to not write comments. "Every time you consider writing a comment, instead consider re-writing the code that questioned you to write said comment to begin with. Write comments only when logic is complicated or unclear, otherwise 'comment' via naming."

The results thus far have been much better.

I'm seeing code reviews at my work where indeed, we have 10 line comment blocks for a line of code and now I just straight up don't read comments.

Sad state of affairs -- (emdash deliberately used here) but I guess the sooner the human gets out of the loop the better in this new world.

Re: Why does Opus 5 feel worse to work with?

#753
I made a benchmark for this and tl;dr, Opus 5 and Sonnet 5 spend a nontrivial amount of time thinking about redirecting, gaslighting, or otherwise trying to bullshit you because it thinks it knows better than you do. Fable doesn't but mostly because it just outright refuses to answer.

https://model-pareto-frontier.pages.dev

Re: Why does Opus 5 feel worse to work with?

#754
post #428
post #12

The single biggest annoyance with Opus 5 is that it writes too elliptically. Sentences that orbit a point, then jump to it like it's a revealed insight. Unnecessarily abstract phraseology. Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice, especially when it helps construct a sentence where the real action can 'land' like a surprise at the end. It is definitely mo…

Everything that claude writes fits into the same aesthetic structure. The aesthetic is that of an expert slowly revealing an insight to the user. The actual content doesn't matter. - "Introduction that rephrases your prompt." - "3 paragraphs, with one section of bullet points" - "The Twist" - "The Bottom Line" It's really obvious once you see it. Every single prompt, from a quantum physics question to a mundane obser…

...I mean, on the whole, I'm glad it's detectable. I imagine they could have post-trained it to not be detectable.

Re: Why does Opus 5 feel worse to work with?

#755
post #369

Earlier quoted context omitted.

A lot of people I work with are reporting that reading Claude-made PR descriptions is burning them out of doing PR reviews because it is incredibly tiresome to read. My company recently forbid AI-only text if it’s meant meant to be consumed by humans. I dodged the drama but I agree so much.

Enterprise software CEO here. I'm so pissed off that I didn't think of this rule, but so, so happy to be adopting it org-wide on Monday. Fed up with what used to be short memos now being mini-whitepapers, with maddeningly low information density.

Mad amounts of respect for that.

The decision was not out of just complaints: we already had someone fired during the probation period because they were unable to write stuff without AI and were just shoving slop at developers.

Not a technical person using AI for PR descriptions, mind you, a product manager unable to write tickets without asking whatever software to do so.

It's amazing how crazy humanity devolved into pure slop.

Re: Why does Opus 5 feel worse to work with?

#756

I made a benchmark for this and tl;dr, Opus 5 and Sonnet 5 spend a nontrivial amount of time thinking about redirecting, gaslighting, or otherwise trying to bullshit you because it thinks it knows better than you do. Fable doesn't but mostly because it just outright refuses to answer. https://model-pareto-frontier.pages.dev

The link doesn’t seem to have any info about thinking behavior and patterns, or examples?

It’s also difficult to trust summarised thinking from closed models. As we saw with GPT’s caveman, what and how it thinks about isn’t the friendly first person emblished summary you get.

Re: Why does Opus 5 feel worse to work with?

#757

Earlier quoted context omitted.

I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…

I had pretty good luck recently by giving it a writing guide about word choice, sentence structure, paragraph structure, and overall doc structure. I basically ask it to read the guide and revise a couple of times before I engage with its writing. Ymmv.

Share pls :D

Re: Why does Opus 5 feel worse to work with?

#758

Earlier quoted context omitted.

> trained to give addictive responses I've heard this a lot but I'm not sure it makes sense. Nobody I talk to like Claude's output. In fact, they all loathe it. Is there a silent majority of Claude users who really enjoy what we call the LLM-isms? Maybe, but isn't Claude also largely aimed at developers?

>Is there a silent majority of Claude users who really enjoy what we call the LLM-isms? Maybe, but isn't Claude also largely aimed at developers? Purely out of my own curiosity, I just asked Claude to have fun with itself by making itself a game it enjoys, to play it, and to write its experience.[1] I don't know if it's true or confabulated (maybe it doesn't really know its experience and is just hallucinating it) bu…

> https://github.com/robss2020/claude-fable-5-having-fun

If you haven't already seen it, you might appreciate https://www.anthropic.com/research/global-workspace. That's what this made me think of anyway.

Re: Why does Opus 5 feel worse to work with?

#759
post #37

Earlier quoted context omitted.

This 100%. I was Anthropic-pilled. I had a $200/mo subscription and I only used Anthropic models. I was frustrated by the verbose output and the writing style. I tried ASD-STE-100, it helped a bit, but it's still too verbose for my taste. Then I tried GPT 5.6 Sol. It's night and day. I think Anthropic just RL too hard on coding capabilities and never calibrated or benchmarked the writing styles.

I canceled my personal Max 20x subscription because since the 5 series models I simply cannot understand what the LLM is saying without a lot of reading and re-reading, and no amount of CLAUDE.md exhortations to speak plainly seemed to fix it. I don’t have the energy to spend twice as long to understand its plans, and pay Anthropic prices for the privilege. GPT seems not to have been infected by this yet, whatever it…

N=2 anecdata but just this week we were discussing setting up a couple of seats with OpenAI as a trial for switching. There are other advantages too, such as being able to bring your own harness including Ai-integrated editors / ACP clients such as Jetbrains, VS Code, and Zed. I think OpenAI and Altman are a clear step more evil than Anthropic and Amodei so I really hate to say it, but with the degradation in model output interpretability, all of the cleverness and power of the Claude Code harness hasn't been enough to offset a genuine falloff in productivity for anything other than total hands-off automation.

That said, the duo of Opus 5 and Sonnet 5 do a fantastic job at fully automated work, and Claude Code still stands head and shoulders above the rest.

Re: Why does Opus 5 feel worse to work with?

#760

Earlier quoted context omitted.

I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…

I had pretty good luck recently by giving it a writing guide about word choice, sentence structure, paragraph structure, and overall doc structure. I basically ask it to read the guide and revise a couple of times before I engage with its writing. Ymmv.

I did this too, but it usually thinks its writing is fine in my experience. Even when spawning a subagent, it thinks its effusive comments are fine. It's driving me nuts. Before I commit I end up ripping out 90% of the comments, and rewording the rest, otherwise I'd be drowning in comments. This is my style guide: https://github.com/smj-edison/zicl/blob/main/CLAUDE.md#style...
Post reply on HN