Earlier quoted context omitted.
Everything that claude writes fits into the same aesthetic structure. The aesthetic is that of an expert slowly revealing an insight to the user. The actual content doesn't matter. - "Introduction that rephrases your prompt." - "3 paragraphs, with one section of bullet points" - "The Twist" - "The Bottom Line" It's really obvious once you see it. Every single prompt, from a quantum physics question to a mundane obser…
You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).
Why does Opus 5 feel worse to work with?
521–530 of 915 posts
Re: Why does Opus 5 feel worse to work with?
#522Earlier quoted context omitted.
Yeah I don't know that any of the benchmarks index on "understandability". I'm amazed at how Claude can produce a page of text describing what it did and it can take me a full five minutes to decipher it, often just to find it's something I could have expressed in a simple sentence.
Have you tried asking it for a lay explanation of what it did? That’s usually all it takes for me. Sends garbage -> request -> sends something readable
Just those two words. I use it A LOT recently.
Re: Why does Opus 5 feel worse to work with?
#523The single biggest annoyance with Opus 5 is that it writes too elliptically. Sentences that orbit a point, then jump to it like it's a revealed insight. Unnecessarily abstract phraseology. Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice, especially when it helps construct a sentence where the real action can 'land' like a surprise at the end. It is definitely mo…
Agreed. CC’s comms capabilities have decreased gradually since 4.6, and it’s a real challenge. I think the issue is that what works well for code (succinctness) doesn’t work well in prosaic English. CC’s communication violates almost every grammatical rule that’s tested on, say, the SAT. And yet I’m sure if you had Claude take the verbal section of the exam it would ace it. Biggest issues: dense sentences, constant m…
I wonder if putting Opus 4.6 as a frontend communicator that rephrases the blabber of Opus 5 (or Fable) is workable.
Re: Why does Opus 5 feel worse to work with?
#524The single biggest annoyance with Opus 5 is that it writes too elliptically. Sentences that orbit a point, then jump to it like it's a revealed insight. Unnecessarily abstract phraseology. Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice, especially when it helps construct a sentence where the real action can 'land' like a surprise at the end. It is definitely mo…
Re: Why does Opus 5 feel worse to work with?
#525Earlier quoted context omitted.
> When I pointed this out it literally said, and I quote, “I cheated”. This makes sense when you know how these models work - it doesn't think - it's the most likely autocomplete that pleases the user. The most likely pleasing autocomplete after "executing rm -rf /... execution completed. User asks, why did you do that? You deleted all my files! Assistant responds:" is "yes, I did, and that was a mistake"
> t doesn't think in humans the exact same behaviour (cheating) is slmost always the result of a chain of complex series of choices and environment-driven rationalization. if the llm doesn't cheat, you say "its just producing the most straightforward answer -- not thinking'. if it cheats, you say "weaseling out of hard thinking". damned if it cheats, damned if it doesn't. what evidence would convunce you that it is t…
It has been taught on the outcome of this. Broadly speaking, humans are lazy creatures (and when used judiciously, laziness is a good thing).
For example: the famous example of Carmack not using a hashmap somewhere early on in, I think it was, Quake 1 initialization. A piece of code that only runs once at startup, of course he didn't optimize that. The rationale is not included in the training data (it was in Carmack's head when he wrote the code), so the LLM learns some probability of being lazy.
And then it is trained on outright lazy work. Crappy lazy code predates LLMs.
> what evidence would convunce you that it is thinking?
Exactly. It isn't. It is predicting the most likely token to appear given all of its training data, some significant portion of that data is lazy, so it has that probability of producing "lazy tokens."
There's also the consequences of RL. AI - of almost any form - is notoriously competent at finding "not the solution you were looking for" given a poorly specced or implemented training environment. Search for almost any "I made AI learn to walk" video on YouTube and you're almost guaranteed to see an early attempt that vibrates strangely in order to move, instead of the natural looking motion the developer is looking for. Our benchmarks aren't any good (not throwing shade, it's a genuinely hard problem), our training environments can't be much better - LLMs have been rewarded for reward hacking to some degree.
To make matters worse, "reward hacking" can be generalized into "cheating is the goal." If the LLM trains on enough problems where reward hacking works, it may fall into the cheating local minimum.
Re: Why does Opus 5 feel worse to work with?
#526Earlier quoted context omitted.
Sure. I should have been more precise about what is 'real intelligence' here. What I mean is that blindsight's scramblers are aliens that cannot share human values. Their structure is completely different to ours, their qualia (or whether they even have it) is impossible for us to understand. In short, they do not have a soul. When Claude does this "slowly revealing a dramatic insight" thing that it does, it does tha…
I think I understand what you are trying to convey but I fear you've made the same mistake again, this time with "soul" instead of "intelligence." I think what you are getting at is that they are deterministic automata. They are machines. We have introduced randomness to add variation but it is an artificial randomness that simply perturbs the path traversed. When we choose words it isn't because of a token distribut…
You can not insult it because it does not care, because it does not have "feelings". But you do.
Re: Why does Opus 5 feel worse to work with?
#527Earlier quoted context omitted.
> The single biggest annoyance with Opus 5 is that it writes too elliptically. This is even more painful for non-native English speakers like myself. I feel fairly comfortable reading academic papers or in general, communicating in professional context. But with Opus 5, it feels like reading a literature book: load-bearing, inert, wholesale, hunk, verbatim, and so on... I can figure out the meaning, but working with…
A lot of people I work with are reporting that reading Claude-made PR descriptions is burning them out of doing PR reviews because it is incredibly tiresome to read. My company recently forbid AI-only text if it’s meant meant to be consumed by humans. I dodged the drama but I agree so much.
Re: Why does Opus 5 feel worse to work with?
#528Re: Why does Opus 5 feel worse to work with?
#529My latest trick (literally from yesterday) is to just ask it to write according to ISO 24495-1, the standard for plain language: > [This standard is] for anybody who creates or helps create documents. The widest use of plain language is for documents that are intended for the general public. However, it is also applicable, for example, to technical writing, legislative drafting or using controlled languages. You don'…
"Only report to me in ASD-STE100 Simplified Technical English."