Live data from Hacker News

Why does Opus 5 feel worse to work with?

mun-logadan.github.io

651–660 of 916 posts

Re: Why does Opus 5 feel worse to work with?

#651

Earlier quoted context omitted.

No, it does not.

You are vastly overestimating the average writer’s ability.

Even then, I don't care about the "average writer". I want great output. I like to imagine that developers have some self-respect, but by now everyone in the industry is spending hundreds of hours every month reading some of the most poorly written prose we could imagine, simply because it affords us to think less.

Re: Why does Opus 5 feel worse to work with?

#652

Earlier quoted context omitted.

I think I understand what you are trying to convey but I fear you've made the same mistake again, this time with "soul" instead of "intelligence." I think what you are getting at is that they are deterministic automata. They are machines. We have introduced randomness to add variation but it is an artificial randomness that simply perturbs the path traversed. When we choose words it isn't because of a token distribut…

An observation: You can never insult an LLM, but it can certainly insult you. You can not insult it because it does not care, because it does not have "feelings". But you do.

Nitpick: The word you meant to use is "offend", not "insult". Just because you can't cause offense to a toaster doesn't mean you can't insult it.

Re: Why does Opus 5 feel worse to work with?

#653

Anthropic, if you're listening - by the time this crops up on Reddit, the front page of HN, etc.... you should be expecting calls from CEOs of major corporations next threatening to abandon ship... We've seen this pattern before several times.. I hope they are listening and address this publicly. I'm not sure what is going on, some users report it works fine or great, others report the degradation. I've experienced b…

You are expecting consistent QoS from a randomly sampled mathematical function.

I get consistency out of the ridiculous pile of quantum noise that is my CPU. Plenty of random processes produce consistency when handled properly. An LLM won't usually give identical outputs for identical inputs, but it's entirely reasonable to expect similar output for similar input when considered on broad metrics like "intelligence" or flowery language or staying on task.

Re: Why does Opus 5 feel worse to work with?

#654
post #650

I really don't get why people think Opus 5 is bad. In my testing it's been fine, but every other model is converging on also being fine. I have ADHD mode installed in my main Claude Code instance though, so that may be part of why I have a better time with it?

What in the world is ADHD mode? Searching it up I see this thing called an “ADHD” skill? Does that work well?

It seems like the skill has some more specific scaffolding for problem solving, so (if that’s true, i didn’t read very in depth) in that case that alone might significantly reduce perfeived performance variability between models

Re: Why does Opus 5 feel worse to work with?

#655

Claude has essentially become useless for agentic development or research. Doesn't matter what model you use. A few rounds and bam, you've burned through your quota. Doesn't matter how "intelligent" their models are, if you can't use them. That, and the quality of AI responses are, in my opinion, significantly worse than competitors like OpenAI. At this pace, I foresee Anthropic becoming the next Nokia. If you would'…

What are you guys doing to burn through limits? I have some dev + prod bots and according to ccusage, use the equivalent of $2500/month with them on CC yet I never hit the rate limits. I feel like I'm using them all the time so I'm curious what you are actually doing that's burning all of these tokens. Can you give me an example? For me, it's: 1. Write a spec for 2. Add design for issue 3. Write code 4. Deploy code a…

I suspect a part of the issue might just be as simple as:

fear of losing context from compaction/starting new chat

then greedy trying to extend/squeeze out answers from the current chat

and being extremely not careful with this just blows through your limits

Re: Why does Opus 5 feel worse to work with?

#656
post #12

The single biggest annoyance with Opus 5 is that it writes too elliptically. Sentences that orbit a point, then jump to it like it's a revealed insight. Unnecessarily abstract phraseology. Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice, especially when it helps construct a sentence where the real action can 'land' like a surprise at the end. It is definitely mo…

Agreed. CC’s comms capabilities have decreased gradually since 4.6, and it’s a real challenge. I think the issue is that what works well for code (succinctness) doesn’t work well in prosaic English. CC’s communication violates almost every grammatical rule that’s tested on, say, the SAT. And yet I’m sure if you had Claude take the verbal section of the exam it would ace it. Biggest issues: dense sentences, constant m…

I’ve been running into this too. It’s especially frustrating when you ask Claude to explain one of its own terms or summaries, and instead of just defining it plainly, it sometimes goes through several rounds of tool calls before giving you a usable explanation. I really don't think such time/tokens should be wasted.

Re: Why does Opus 5 feel worse to work with?

#657
post #610

Earlier quoted context omitted.

You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).

Alas, it writes so much better than the average human that it's what everyone started using. Hence the utter familiarity and now contempt.

I agree that it’s better at writing than a 50%-ile human, but it’s worse at communicating through writing than most humans.

Even an average human writer can communicate details much more succinctly and directly than an LLM

Re: Why does Opus 5 feel worse to work with?

#658
post #491
post #428

Earlier quoted context omitted.

Everything that claude writes fits into the same aesthetic structure. The aesthetic is that of an expert slowly revealing an insight to the user. The actual content doesn't matter. - "Introduction that rephrases your prompt." - "3 paragraphs, with one section of bullet points" - "The Twist" - "The Bottom Line" It's really obvious once you see it. Every single prompt, from a quantum physics question to a mundane obser…

It's the same way how every AI generated poster looks exactly the same. As if there is a single underlying prompt that describes the template of the poster/long-form article, and it does not dare deviate from that.

Isn't there? Like everyone using $MODEL is starting from the same base system-prompt. Then our user input is a small bit on top of that core mode. Like what would happen if everyone asked Mikey to paint their ceiling - they'd all be similar and therefore boring.

Re: Why does Opus 5 feel worse to work with?

#659
I'm glad this is being talked about. I noticed it too. I find Opus 5 to be overly (and unhelpfully) critical, in a sort of well-actually way. It ignores nuance in my direction or prompts.

It is better at engineering tasks; I've seen an appreciable difference in its problem-solving abilities. But perhaps that same thing makes it kind of an annoying prick to work with on anything non-engineering, for which I stick to 4.8, where the prose is a little more florid rather than pugnacious.

Re: Why does Opus 5 feel worse to work with?

#660

Earlier quoted context omitted.

I'll say everything indicates we've hit or are near peak for the masses at least (unless you start paying 50x more) but to each his own. 4.6 was best for us and right now yeah OpenAI and others are edging forward, but slower while prices are increasing industry wide as much as 20x, time to completion is increasing wildly and i'm sure they'll do the same over at OpenAI as their compute constraints also start to take a…

Literally nothing indicates we've hit a peak, but I guess I'm discussing this with someone who thinks every iteration since Opus 4.6 had zero ROI so there's probably not much common ground here.

I think you misunderstand what i'm saying: the companies are not profitable yet, it's the ROI on the investments, not that these tools are useless, very much the opposite, but the business model is not viable, hence the price increase and degrading quality, slower responses etc. I agree theres still progress but its slowing and we're probably near peak.
Post reply on HN