Live data from Hacker News

Why does Opus 5 feel worse to work with?

mun-logadan.github.io

641–650 of 916 posts

Re: Why does Opus 5 feel worse to work with?

#641
post #428

Earlier quoted context omitted.

Everything that claude writes fits into the same aesthetic structure. The aesthetic is that of an expert slowly revealing an insight to the user. The actual content doesn't matter. - "Introduction that rephrases your prompt." - "3 paragraphs, with one section of bullet points" - "The Twist" - "The Bottom Line" It's really obvious once you see it. Every single prompt, from a quantum physics question to a mundane obser…

You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).

This is painfully accurate.

Re: Why does Opus 5 feel worse to work with?

#642
post #28

Earlier quoted context omitted.

From the little i understand that wouldnt be an issue because the model is ‘just’ using interchangeable words in a mathematical non-random way. Like using the same number of adjectives and the exct same words, but in a order that wouldn’t be mathematically plausible unless it was the watermark

I wouldn’t exactly put it like that. It’s moreso the model sometimes outputting non-optimal tokens in a way that’s detectable if you know the algorithm. It seems possible for that to make the response “drift” far from what it would’ve been, because it’s constant entropy that adds up after time. (However, according to Anthropic and Google, it doesn’t really impact the quality of responses. I find that a bit hard to be…

What next token is “optimal” is fuzzy and subjective. All transformer based models have a “temperature” setting whose sole purpose is to randomly make choices other than the most likely next token. This is crucial to good output, but you wouldn’t call those choices “non-optimal” even if they are less likely. In any text generation task there are constant opportunities to make a choice from equivalent options.

Re: Why does Opus 5 feel worse to work with?

#643
post #12

The single biggest annoyance with Opus 5 is that it writes too elliptically. Sentences that orbit a point, then jump to it like it's a revealed insight. Unnecessarily abstract phraseology. Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice, especially when it helps construct a sentence where the real action can 'land' like a surprise at the end. It is definitely mo…

This is what news headlines did for decades to bait you into reading the details. I wouldn't be surprised if AI companies do that intentionally to consume more tokens trying to understand what had just happened

Re: Why does Opus 5 feel worse to work with?

#644
post #120

I’ve been doing some heavy work on a personal project lately. I burned through the limits on Claude, the plus a few hundred dollars in credits, and ultimately decided to move to an OpenAI account just so I can keep going. I was surprised to find that OpenAI Sol is much much nicer to work with than Opus 5 or Fable at the moment. Especially on Opus 5, the way it communicates is just exhausting. It keeps “being honest”…

I’ve been doing OCR of scans of old magazines (specifically extracting music reviews and charts, so turning complex layouts into structured data) and have been impressed with GPT/Codex’s performance.

My setup has a Sol orchestrator and Terra OCR agents and seems to get great results. I’ve not dug into the details too much, it also has a Tesseract stage as an deterministic input which it told me helped. Not sure how token efficient it is but I often don’t have anything to do with my personal tokens ahead of a reset so just let it burn through it in batches.

I am impressed (both in this task and other work I’ve done) not just at how well Codex can setup a structure for a complex task like this, but how it will keep going (Claude seems to find excuses to stop) and also can critique and refine its approach as it goes.

I did try out a bunch of other models and specific OCR providers but none of them hit the same accuracy for my task as Codex so I’m sticking with it.

Re: Why does Opus 5 feel worse to work with?

#645

I've gone back to 4.8. 5 would constantly veer of in random directions if not working from 100% strict and narrow instructions. I find it weird there's not more discussion here on HN on how the most used model now has clearly degraded in quality and it seems we've hit a peak and are on a downslope - because the model is clearly smaller or more economical for Anthropic no doubt about it, and the benchmaxxing they do i…

I haven't had time to complain on here because I spend all my time trying to work out wtf Opus 5 is talking about and googling words I've never seen or heard in 40 years as a native English speaker.

I’m either out of touch with contemporary terminology but Opus 5 dropped “pre-mortem” on me today. Figured I’d just figure it out with more context.

Re: Why does Opus 5 feel worse to work with?

#646
post #600

Earlier quoted context omitted.

I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…

> if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it There are sociological reasons why this happens less with humans: 1. You cycle your dumb repetitive jokes with everyone you meet, so nobody hears it twice 2. Those who know you well will notice when you're just repeating ("dad jokes") 3. As a person's idiosyncrasies are beginning to…

>Perhaps between model version releases, frontier labs can harvest the web and ask "What Claudisms do people mention negatively?" but I don't think they do that yet.

That won't happen. People can't phrase their objections in a succinct-enough way. When they do, the objection is superficial ("too many em dashes") and doesn't strike at the core of what makes LLM output bad.

Re: Why does Opus 5 feel worse to work with?

#647

Earlier quoted context omitted.

I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…

> I am curious why LLM writing has such an uncanny valley feel to it. Because they are HEAVILY trained to give addictive responses. They don't want to just answer your question. They want to sycophantically make you feel like a genius for being smart enough to use them.

> trained to give addictive responses

I've heard this a lot but I'm not sure it makes sense. Nobody I talk to like Claude's output. In fact, they all loathe it.

Is there a silent majority of Claude users who really enjoy what we call the LLM-isms? Maybe, but isn't Claude also largely aimed at developers?

Re: Why does Opus 5 feel worse to work with?

#648
I'm not sure if it's worse, as much as it seems different, to a different degree.

Each model update changes how to best prompt with it, since that's the words that are used with it generically or specifically it can hit some people, and not others, or more, and not less.

Re: Why does Opus 5 feel worse to work with?

#649

A gem Opus 5 gifted to me today: "A devastating pair of findings, and the first is beautiful in a way worth naming: the anti-vacuity floor is what blinds the gate to a vacuous case."

Oh yes, Claude seems to be loving the word vacuous recently. My test bed side project is full of vacuous this and that now.

I even try and get it to define what it classifies as vacuous and it can’t do so without getting stuck in some kind of trap. It’s like a word with some kind of huge gravity for it.

Re: Why does Opus 5 feel worse to work with?

#650
I really don't get why people think Opus 5 is bad. In my testing it's been fine, but every other model is converging on also being fine. I have ADHD mode installed in my main Claude Code instance though, so that may be part of why I have a better time with it?
Post reply on HN