Earlier quoted context omitted.
Everything that claude writes fits into the same aesthetic structure. The aesthetic is that of an expert slowly revealing an insight to the user. The actual content doesn't matter. - "Introduction that rephrases your prompt." - "3 paragraphs, with one section of bullet points" - "The Twist" - "The Bottom Line" It's really obvious once you see it. Every single prompt, from a quantum physics question to a mundane obser…
> The aesthetic is that of an expert slowly revealing an insight to the user. Ah, that's it! Thank you. I wonder if they are training it to talk like this because that's what their customers actually want? They want a machine genius to lead them.
Why does Opus 5 feel worse to work with?
761–770 of 916 posts
Re: Why does Opus 5 feel worse to work with?
#762Earlier quoted context omitted.
It’s the repetitiveness of style, the attempt to make everything seem as impactful as possible, the use of short sentences (too much Hemingway in the training data?), and obvious patterns like “it’s not this, it’s that” and several others. Real human writing doesn’t follow such strict rules. When the same small set of rules is applied over and over throughout a text, it becomes obviously strange and machine-like.
I blame RLHF entirely for this. Nobody used to talk like AI speech before.
Re: Why does Opus 5 feel worse to work with?
#763Re: Why does Opus 5 feel worse to work with?
#764Earlier quoted context omitted.
Interestingly, I've been noticing almost the opposite issue. 5.6 Sol frequently starts its responses with "Yes" even when my prompt doesn't contain a yes-or-no question.
I've been noticing the same thing since 5.4- it starts with "Yes" yet I haven't asked a question. I think it might be related to the reasoning, like it's answering its own questions?
The "It's not X, it's Y"-style repetition is another example of that: It argues with itself in the background, and the argument leaks out as if in refutation of things that you (the human) never had imparted into discussion at all.
Re: Why does Opus 5 feel worse to work with?
#765A gem Opus 5 gifted to me today: "A devastating pair of findings, and the first is beautiful in a way worth naming: the anti-vacuity floor is what blinds the gate to a vacuous case."
Certainly a human would write something much clearer than yours. Maybe: "Two good findings here. #1: There's already a min value on the gate and anything lower doesn't pass through, so we don't need special handling for the zero case."
Re: Why does Opus 5 feel worse to work with?
#766Re: Why does Opus 5 feel worse to work with?
#767Earlier quoted context omitted.
You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).
I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…
This is the thing that LLM writing doesn't really do. Since there isn't a mental state, there's nothing to mirror anyway, and somehow you can detect the absence of it even though it's hard to put any of the machinery into words. Corporate speech also fails to activate this machinery in the same way, but LLMs seem to do it more egregiously, probably because the corporate speech was at least compiled by a human---even if it is not the thoughts of an individual, it is the "thoughts" of an "entity", the abstract corporation, which the writer was speaking for, and so you can still wrap your mind around the fact that it is communicating with you.
Re: Why does Opus 5 feel worse to work with?
#768Me: "Review the following and work through the implementation"
Opus 5: "Called tool , Called tool ..." - for a few minutes.
Opus 5: "I've implemented X, do you want me to commit changes?"
Me: "None of those changes are on the file system"
Opus 5: "You're right, all the tool calls were fabricated."
Re: Why does Opus 5 feel worse to work with?
#769Earlier quoted context omitted.
May be getting harder to catch the mistakes but that makes them worse in my view. I'd much rather they make easy to spot mistakes because I don't expect factuality from them anyway just speedy transformation of information I already have available. In fact it's the lossiest transformation tool I've ever used and it's still useful despite that. If it reaches one nine of reliability that would be huge but given the pac…
Or, the hard to catch mistakes were always there, and now we focus on them instead of the obvious ones that have been eliminated.
The number of mistakes are also about the same.
But more mistakes are harder to catch. The output is more polished and convoluted and that make errors harder to catch. I don't want that.
Re: Why does Opus 5 feel worse to work with?
#770Earlier quoted context omitted.
That is very strange. I havent seen that. It sounds like something leaking from its system prompt or something that its not handling well. Anthropic trying to prevent it from running longer or something.
I definitely have. It's like something in the system prompt has "keep the well-being of the human in mind" and of course the date and time, but something makes the model take that way way too literally.
Whenever context gets towards the max length is when I've noticed it.