Earlier quoted context omitted.
I had this debate with my coworker who prefers anthropic models to open ai ones. I ended up settling into the idea that gpt 5.6 is better used as a tool and opus 5 is a companion. GPT 5.6 takes you literally whereas opus 5 tends to take more liberties to try to get to the “spirit” of what you want. It comes down to preference, and I don’t want a companion.
I've switched to Codex a few months ago when Claude's weekly limits were getting pretty stiff, and I haven't looked back. Both GPT 5.5 and 5.6 are quite capable, especially compared to nerfed Opus 4.7 (haven't tried 5). Also, the Codex guy regularly resets weekly limits for everyone, which is a nice bonus (I know it's a temporary gimmick to attract more users, but I might as well use it while it lasts.)
Why does Opus 5 feel worse to work with?
701–710 of 915 posts
Re: Why does Opus 5 feel worse to work with?
#702Since they have become so capable the new bottleneck is what they can't know. The stuff that's inside people's brains who work in real companies with products and processes absent from any training set.
Re: Why does Opus 5 feel worse to work with?
#703A gem Opus 5 gifted to me today: "A devastating pair of findings, and the first is beautiful in a way worth naming: the anti-vacuity floor is what blinds the gate to a vacuous case."
Re: Why does Opus 5 feel worse to work with?
#704Earlier quoted context omitted.
> I am curious why LLM writing has such an uncanny valley feel to it. Because they are HEAVILY trained to give addictive responses. They don't want to just answer your question. They want to sycophantically make you feel like a genius for being smart enough to use them.
> trained to give addictive responses I've heard this a lot but I'm not sure it makes sense. Nobody I talk to like Claude's output. In fact, they all loathe it. Is there a silent majority of Claude users who really enjoy what we call the LLM-isms? Maybe, but isn't Claude also largely aimed at developers?
Re: Why does Opus 5 feel worse to work with?
#705Earlier quoted context omitted.
You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).
I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…
Re: Why does Opus 5 feel worse to work with?
#706Earlier quoted context omitted.
I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…
> if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it There are sociological reasons why this happens less with humans: 1. You cycle your dumb repetitive jokes with everyone you meet, so nobody hears it twice 2. Those who know you well will notice when you're just repeating ("dad jokes") 3. As a person's idiosyncrasies are beginning to…
Re: Why does Opus 5 feel worse to work with?
#707Earlier quoted context omitted.
You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).
I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…
Re: Why does Opus 5 feel worse to work with?
#708Earlier quoted context omitted.
>Perhaps between model version releases, frontier labs can harvest the web and ask "What Claudisms do people mention negatively?" but I don't think they do that yet. That won't happen. People can't phrase their objections in a succinct-enough way. When they do, the objection is superficial ("too many em dashes") and doesn't strike at the core of what makes LLM output bad.
Don't talk about other people in such a way. If we had a way in Claude to mark a word/token and downvote/reduce it logits then it wouldn't be to hard to send loadbearing to token valhalla. In the future, I hope we get a way to randomize the language idiosyncrasies and/or personalities better.
Re: Why does Opus 5 feel worse to work with?
#709Re: Why does Opus 5 feel worse to work with?
#710Earlier quoted context omitted.
> t doesn't think in humans the exact same behaviour (cheating) is slmost always the result of a chain of complex series of choices and environment-driven rationalization. if the llm doesn't cheat, you say "its just producing the most straightforward answer -- not thinking'. if it cheats, you say "weaseling out of hard thinking". damned if it cheats, damned if it doesn't. what evidence would convunce you that it is t…
If I could give it a novel task outside of its explicit training and see it actually improve just through accreting context, I'd be convinced it was thinking. The opposite happens in practice. I test new models with two tasks: iteratively generating SVGs based on a text description with rendered rasters for feedback; and generating "Before and After" clues like on Jeopardy, where the response has two overlapping phra…