Live data from Hacker News

Why does Opus 5 feel worse to work with?

mun-logadan.github.io

601–610 of 916 posts

Re: Why does Opus 5 feel worse to work with?

#601
post #428

Earlier quoted context omitted.

Everything that claude writes fits into the same aesthetic structure. The aesthetic is that of an expert slowly revealing an insight to the user. The actual content doesn't matter. - "Introduction that rephrases your prompt." - "3 paragraphs, with one section of bullet points" - "The Twist" - "The Bottom Line" It's really obvious once you see it. Every single prompt, from a quantum physics question to a mundane obser…

Perhaps this is related to their new "invisible watermark" concept which would probably require rather contrived language patterns to make possible.

I love this theory. "We've invented a new invisible watermark that can detect whether code is LLM written."

The watermark: counting instances of 'load-bearing seam', 'the hard truth', 'and that's the whole point'.

Re: Why does Opus 5 feel worse to work with?

#602

Earlier quoted context omitted.

I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…

> I am curious why LLM writing has such an uncanny valley feel to it. Because they are HEAVILY trained to give addictive responses. They don't want to just answer your question. They want to sycophantically make you feel like a genius for being smart enough to use them.

This makes some good intuitive sense, but to me the sycophancy feels like it is an emergent property of turning a next word predictor into a conversational chatbot whether or not it’s intentionally trained that way. Your prompt and its earlier responses is all it has in its context window, so of course it lends undue importance to everything you say. Does that seem like a contributing factor to you?

Re: Why does Opus 5 feel worse to work with?

#603
post #545

Earlier quoted context omitted.

That’s how I see it too. Claude is more “fun” to use, like a coworker I have to talk to now and then to steer it, while gpt-5.6 is a task machine: I give it a task and it is very consistent, reliable and predictable in its execution. I don’t have to interrupt it, it gets the task done exactly how I wanted it, but it’s “boring” and feels more sterile

Reminds me of this: https://www.geoffreylitt.com/2025/07/27/enough-ai-copilots-w... I think this is such a great reframing. It makes so much sense; I need an AI that acts more as a HUD and gives me superpowers, not just a copilot that can tell me when I've misspelled a word.

False dichotomy, no? You can have a HUD, and a copilot, and your copilot can also have a HUD. And to complete the idea, you can also have neither.

Re: Why does Opus 5 feel worse to work with?

#604
post #547
post #187

Earlier quoted context omitted.

I notice the models with reasoning can conflate “internal” (or subagent) discussions with external (i.e. me). So it is accurately indicating “I’ve had this discussion before” but incorrectly asserting who it was with. My understanding of how “thinking”works is limited though, and given the reduced visibility into the thinking traces, it is harder to tell if this is actually happening or if these are imaginary discuss…

Yes. I've not used Opus 5 much directly, but when it was Fable and Opus 4.8, I found Fable did this all the time and it was maddening. It'd say stuff like "Oh, I mentioned that between tool calls" or something.

I’m pretty sure this is a Claude code bug - if you do ctrl+o you can see those hidden responses from Fable. Fable doesn’t know the harness is bugged, so I added instruction to my Claude.md to save all commentary for final message.

Re: Why does Opus 5 feel worse to work with?

#605
post #187

Earlier quoted context omitted.

I notice the models with reasoning can conflate “internal” (or subagent) discussions with external (i.e. me). So it is accurately indicating “I’ve had this discussion before” but incorrectly asserting who it was with. My understanding of how “thinking”works is limited though, and given the reduced visibility into the thinking traces, it is harder to tell if this is actually happening or if these are imaginary discuss…

Oh, that's interesting - because that's absolutely what's happening in my experience. If I look at the thinking (which seems to have become unavailable in Opus 5 a lot of the time, but was present - and often useful - in 4.8/4.6) you're right - it's having the discussion with itself, and seems unable to distinguish that discussion from discussions with me. BUT it also seems to be related to the length of the chat - t…

With GPT 5.6 Luna the thinking once or twice leaked into the output for me. It's interesting, but perhaps not particularly useful.

It would be endless paragraphs of something among the lines of:

Need prepare final response? Yes provide. But wait, chat tool complete? Final needed but user already complete. Need summary, preparing final. Response complete. Wait but is final response complete? Need provide. Start finalizing now but wait did user acknowledge final complete? Assistant response final: user complete. Should now create final?

Re: Why does Opus 5 feel worse to work with?

#606

I've gone back to 4.8. 5 would constantly veer of in random directions if not working from 100% strict and narrow instructions. I find it weird there's not more discussion here on HN on how the most used model now has clearly degraded in quality and it seems we've hit a peak and are on a downslope - because the model is clearly smaller or more economical for Anthropic no doubt about it, and the benchmaxxing they do i…

I haven't had time to complain on here because I spend all my time trying to work out wtf Opus 5 is talking about and googling words I've never seen or heard in 40 years as a native English speaker.

Re: Why does Opus 5 feel worse to work with?

#607

Earlier quoted context omitted.

You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).

I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…

[dead]

Re: Why does Opus 5 feel worse to work with?

#609

Earlier quoted context omitted.

You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).

I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…

Might be related to their fingerprinting of llm output they said earlier in the week.

Re: Why does Opus 5 feel worse to work with?

#610
post #428

Earlier quoted context omitted.

Everything that claude writes fits into the same aesthetic structure. The aesthetic is that of an expert slowly revealing an insight to the user. The actual content doesn't matter. - "Introduction that rephrases your prompt." - "3 paragraphs, with one section of bullet points" - "The Twist" - "The Bottom Line" It's really obvious once you see it. Every single prompt, from a quantum physics question to a mundane obser…

You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).

Alas, it writes so much better than the average human that it's what everyone started using. Hence the utter familiarity and now contempt.
Post reply on HN