Earlier quoted context omitted.
No, it does not.
You are vastly overestimating the average writer’s ability.
Why does Opus 5 feel worse to work with?
651–660 of 916 posts
Re: Why does Opus 5 feel worse to work with?
#652Earlier quoted context omitted.
I think I understand what you are trying to convey but I fear you've made the same mistake again, this time with "soul" instead of "intelligence." I think what you are getting at is that they are deterministic automata. They are machines. We have introduced randomness to add variation but it is an artificial randomness that simply perturbs the path traversed. When we choose words it isn't because of a token distribut…
An observation: You can never insult an LLM, but it can certainly insult you. You can not insult it because it does not care, because it does not have "feelings". But you do.
Re: Why does Opus 5 feel worse to work with?
#653Anthropic, if you're listening - by the time this crops up on Reddit, the front page of HN, etc.... you should be expecting calls from CEOs of major corporations next threatening to abandon ship... We've seen this pattern before several times.. I hope they are listening and address this publicly. I'm not sure what is going on, some users report it works fine or great, others report the degradation. I've experienced b…
You are expecting consistent QoS from a randomly sampled mathematical function.
Re: Why does Opus 5 feel worse to work with?
#654I really don't get why people think Opus 5 is bad. In my testing it's been fine, but every other model is converging on also being fine. I have ADHD mode installed in my main Claude Code instance though, so that may be part of why I have a better time with it?
It seems like the skill has some more specific scaffolding for problem solving, so (if that’s true, i didn’t read very in depth) in that case that alone might significantly reduce perfeived performance variability between models
Re: Why does Opus 5 feel worse to work with?
#655Claude has essentially become useless for agentic development or research. Doesn't matter what model you use. A few rounds and bam, you've burned through your quota. Doesn't matter how "intelligent" their models are, if you can't use them. That, and the quality of AI responses are, in my opinion, significantly worse than competitors like OpenAI. At this pace, I foresee Anthropic becoming the next Nokia. If you would'…
What are you guys doing to burn through limits? I have some dev + prod bots and according to ccusage, use the equivalent of $2500/month with them on CC yet I never hit the rate limits. I feel like I'm using them all the time so I'm curious what you are actually doing that's burning all of these tokens. Can you give me an example? For me, it's: 1. Write a spec for 2. Add design for issue 3. Write code 4. Deploy code a…
fear of losing context from compaction/starting new chat
then greedy trying to extend/squeeze out answers from the current chat
and being extremely not careful with this just blows through your limits
Re: Why does Opus 5 feel worse to work with?
#656The single biggest annoyance with Opus 5 is that it writes too elliptically. Sentences that orbit a point, then jump to it like it's a revealed insight. Unnecessarily abstract phraseology. Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice, especially when it helps construct a sentence where the real action can 'land' like a surprise at the end. It is definitely mo…
Agreed. CC’s comms capabilities have decreased gradually since 4.6, and it’s a real challenge. I think the issue is that what works well for code (succinctness) doesn’t work well in prosaic English. CC’s communication violates almost every grammatical rule that’s tested on, say, the SAT. And yet I’m sure if you had Claude take the verbal section of the exam it would ace it. Biggest issues: dense sentences, constant m…
Re: Why does Opus 5 feel worse to work with?
#657Earlier quoted context omitted.
You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).
Alas, it writes so much better than the average human that it's what everyone started using. Hence the utter familiarity and now contempt.
Even an average human writer can communicate details much more succinctly and directly than an LLM
Re: Why does Opus 5 feel worse to work with?
#658Earlier quoted context omitted.
Everything that claude writes fits into the same aesthetic structure. The aesthetic is that of an expert slowly revealing an insight to the user. The actual content doesn't matter. - "Introduction that rephrases your prompt." - "3 paragraphs, with one section of bullet points" - "The Twist" - "The Bottom Line" It's really obvious once you see it. Every single prompt, from a quantum physics question to a mundane obser…
It's the same way how every AI generated poster looks exactly the same. As if there is a single underlying prompt that describes the template of the poster/long-form article, and it does not dare deviate from that.
Re: Why does Opus 5 feel worse to work with?
#659It is better at engineering tasks; I've seen an appreciable difference in its problem-solving abilities. But perhaps that same thing makes it kind of an annoying prick to work with on anything non-engineering, for which I stick to 4.8, where the prose is a little more florid rather than pugnacious.
Re: Why does Opus 5 feel worse to work with?
#660Earlier quoted context omitted.
I'll say everything indicates we've hit or are near peak for the masses at least (unless you start paying 50x more) but to each his own. 4.6 was best for us and right now yeah OpenAI and others are edging forward, but slower while prices are increasing industry wide as much as 20x, time to completion is increasing wildly and i'm sure they'll do the same over at OpenAI as their compute constraints also start to take a…
Literally nothing indicates we've hit a peak, but I guess I'm discussing this with someone who thinks every iteration since Opus 4.6 had zero ROI so there's probably not much common ground here.