Live data from Hacker News

Why does Opus 5 feel worse to work with?

mun-logadan.github.io

761–770 of 916 posts

Re: Why does Opus 5 feel worse to work with?

#761
post #737
post #428

Earlier quoted context omitted.

Everything that claude writes fits into the same aesthetic structure. The aesthetic is that of an expert slowly revealing an insight to the user. The actual content doesn't matter. - "Introduction that rephrases your prompt." - "3 paragraphs, with one section of bullet points" - "The Twist" - "The Bottom Line" It's really obvious once you see it. Every single prompt, from a quantum physics question to a mundane obser…

> The aesthetic is that of an expert slowly revealing an insight to the user. Ah, that's it! Thank you. I wonder if they are training it to talk like this because that's what their customers actually want? They want a machine genius to lead them.

It’s the TED Talk playbook: the crafting of a lecture given by an expert to laypeople to maximize attention, engagement and satisfaction. Every piece of prose is built to pack in as many TED Talk mic-drops/expectation-subverting insight bombs as possible.

Re: Why does Opus 5 feel worse to work with?

#762
post #671

Earlier quoted context omitted.

It’s the repetitiveness of style, the attempt to make everything seem as impactful as possible, the use of short sentences (too much Hemingway in the training data?), and obvious patterns like “it’s not this, it’s that” and several others. Real human writing doesn’t follow such strict rules. When the same small set of rules is applied over and over throughout a text, it becomes obviously strange and machine-like.

I blame RLHF entirely for this. Nobody used to talk like AI speech before.

So is it the LLM or us that's getting the RLHF? /s

Re: Why does Opus 5 feel worse to work with?

#764
post #640

Earlier quoted context omitted.

Interestingly, I've been noticing almost the opposite issue. 5.6 Sol frequently starts its responses with "Yes" even when my prompt doesn't contain a yes-or-no question.

I've been noticing the same thing since 5.4- it starts with "Yes" yet I haven't asked a question. I think it might be related to the reasoning, like it's answering its own questions?

It's been happening for years, and it does seem to be entirely related to its background reasoning leaking out of the context and into the output.

The "It's not X, it's Y"-style repetition is another example of that: It argues with itself in the background, and the argument leaks out as if in refutation of things that you (the human) never had imparted into discussion at all.

Re: Why does Opus 5 feel worse to work with?

#765

A gem Opus 5 gifted to me today: "A devastating pair of findings, and the first is beautiful in a way worth naming: the anti-vacuity floor is what blinds the gate to a vacuous case."

I always have to read these kind of Claude sentences (yours is a great example) two or three times to really understand what they're trying to say, which almost never happens when I read human writing. Sometimes I'm not certain if it's because the AI's writing is super dense with information - or whether it's the opposite and the small core of information is surrounded by a surplus of junk.

Certainly a human would write something much clearer than yours. Maybe: "Two good findings here. #1: There's already a min value on the gate and anything lower doesn't pass through, so we don't need special handling for the zero case."

Re: Why does Opus 5 feel worse to work with?

#766
I don't use Claude for my daily work anymore(due to OAuth restrictions on third party agents), but one theory I saw in another community why Opus 5 is so bad even though benchmark scores were good, is that Anthropic's internal usage pattern is to let Fable to spawn and manage swarms of Opus subagents. This pattern won't penalize that Opus is not well aligned for direct human coworking on posttraining. The worse part is, this makes Fable the user's best default choice for every jobs even if they don't have unlimited credits like Anthropic employees do.

Re: Why does Opus 5 feel worse to work with?

#767

Earlier quoted context omitted.

You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).

I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…

I think that there is a sort of mechanism by which, when a human speaks to you, you can "mirror" their internal mental state, and the quality of writing often corresponds to how much you get pleasure or information or whatever your goal is from that mental state. The important part is that the words themselves are just pointers to the state. So, the person speaking, if they are skillful, gives enough words, and enough variety, that you can start to produce a state yourself which resembles theirs.

This is the thing that LLM writing doesn't really do. Since there isn't a mental state, there's nothing to mirror anyway, and somehow you can detect the absence of it even though it's hard to put any of the machinery into words. Corporate speech also fails to activate this machinery in the same way, but LLMs seem to do it more egregiously, probably because the corporate speech was at least compiled by a human---even if it is not the thoughts of an individual, it is the "thoughts" of an "entity", the abstract corporation, which the writer was speaking for, and so you can still wrap your mind around the fact that it is communicating with you.

Re: Why does Opus 5 feel worse to work with?

#768
Very timely, just this morning:

Me: "Review the following and work through the implementation"

Opus 5: "Called tool , Called tool ..." - for a few minutes.

Opus 5: "I've implemented X, do you want me to commit changes?"

Me: "None of those changes are on the file system"

Opus 5: "You're right, all the tool calls were fabricated."

Re: Why does Opus 5 feel worse to work with?

#769
post #742

Earlier quoted context omitted.

May be getting harder to catch the mistakes but that makes them worse in my view. I'd much rather they make easy to spot mistakes because I don't expect factuality from them anyway just speedy transformation of information I already have available. In fact it's the lossiest transformation tool I've ever used and it's still useful despite that. If it reaches one nine of reliability that would be huge but given the pac…

Or, the hard to catch mistakes were always there, and now we focus on them instead of the obvious ones that have been eliminated.

I still see obvious mistakes, so not eliminated.

The number of mistakes are also about the same.

But more mistakes are harder to catch. The output is more polished and convoluted and that make errors harder to catch. I don't want that.

Re: Why does Opus 5 feel worse to work with?

#770
post #743

Earlier quoted context omitted.

That is very strange. I havent seen that. It sounds like something leaking from its system prompt or something that its not handling well. Anthropic trying to prevent it from running longer or something.

I definitely have. It's like something in the system prompt has "keep the well-being of the human in mind" and of course the date and time, but something makes the model take that way way too literally.

Personally I think it's due to context length.

Whenever context gets towards the max length is when I've noticed it.

Post reply on HN