Live data from Hacker News

Why does Opus 5 feel worse to work with?

mun-logadan.github.io

891–900 of 915 posts

Re: Why does Opus 5 feel worse to work with?

#891
post #12

The single biggest annoyance with Opus 5 is that it writes too elliptically. Sentences that orbit a point, then jump to it like it's a revealed insight. Unnecessarily abstract phraseology. Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice, especially when it helps construct a sentence where the real action can 'land' like a surprise at the end. It is definitely mo…

Pondering this one night last week, I realized that because LLMs can only reason with written language, what we might be seeing emerge with Opus’s load-bearing mumbo jumbo is its own creole for structural reasoning. Not only are our brains wide, our senses are, too. I slow down to a crawl when I have to read actual math in a CS paper, but show me diagrams and I can reason about whatever sort of data structure or algo…

> because LLMs can only reason with written language

Is that right? Think of code that draws a square, the representation of that code in storage, the movement of electrons, and the 'actual' square on the screen -- these are all to us transcriptions of same thing across different domains or media, and we can deterministically translate back and forth between them, but there's no real, actual, essential identity property. The code isn't the square. The electricity isn't the code. The storage isn't the electricity.

I think of LLM reasoning and output similarly. The underlying graph of weights, matrices, and other data aren't knowledge, understanding, or language. 'Translating' the system's output to language is jusas valid and correct as translating it into some visual representation that would be incoherent to us, like a sequence of flashing lights or imperceptible noise patterns projected over an image of a dog.

I guess this is all a very long winded way of restating Chinese-room problem: we feed the man in the room a message; he returns one that, for all the world, is indistinguishable from a "real" response that you and I might send, but, like you said, he has no access to sense data. He also has no access to the biology underlying real mental processes. He also doesn't have any personhood that we can discern. He has, rather, gradually developed through reinforcement the tendency to provide responses approximating all of all of that.

I'm not sure the epistemological question "does he understand" (which is what the Chinese-room problem asks) has any meaningful answer. There's no mind, so there's no understanding. What there is, rather, is a system that generate patterns that we map to language and that our brains therefore map to communication, personhood, meaning, etc. It's the square I mentioned earlier. It's to us a convincing simulacrum, and it may be faithful enough to us to stand in those things, but that's not what it is.

My sense is that LLMs are (a) the big-data Pyramids of Giza and (b) a consequence of hardware and software developing ways of generating abstractions that capture and generate more complex patterns than were previously possible to capture or generate in a manner comprehensible to humans. Everything humans do follows some kind of pattern. Language is the perfect way for a machine to capture that, because is simpler than the world itself and has clear rules and patterns, encoded in representations computers already have, that, modeled well enough, can generate output indistinguishable from the real thing -- what you or I might do with it.

But the pattern matching that it does, and that we translate into language, is much more numerically rigorous and complex than anything you or I consciously do with words (hell, most people can't even figure out when to use "lie" vs "lay") and not doing what you or I do with it. It's not language. It's the square on the screen.

Re: Why does Opus 5 feel worse to work with?

#892

Earlier quoted context omitted.

There's less consensus around what you're talking about that you imagine, which was my initial point. Smart people aren't just this academic bubble that you think it is. Not seeing such views made you develop this association of "smart" has to mean "interprets cultural products according to my political ideology". My reactive reaction was to point out that this isn't so.

I didn't make any claim about being smart or not aside from brief mockery of people who care about the word and things like IQ tests I'd personally trust a rigorous, well-evinced, and methodical dissection of a topic by someone with completely average intelligence over a lazy broad generalization by someone who scored high on a paper test that people (wrongly) assume maps onto the ability to assess and analyze the wo…

I've participated in peer review from both sides, and it's much less impressive than people imagine from the outside.

Re: Why does Opus 5 feel worse to work with?

#893
post #671

Earlier quoted context omitted.

It’s the repetitiveness of style, the attempt to make everything seem as impactful as possible, the use of short sentences (too much Hemingway in the training data?), and obvious patterns like “it’s not this, it’s that” and several others. Real human writing doesn’t follow such strict rules. When the same small set of rules is applied over and over throughout a text, it becomes obviously strange and machine-like.

I blame RLHF entirely for this. Nobody used to talk like AI speech before.

What's weird though is how consistent it is, even across models to some extent. Were the RLHF people given a really specific style guide?

While I agree that no-one used to write like that as a whole before, all the elements can be found in different places. Short sentences to avoid discouraging poor readers. Maximally impactful statements are commonly used in marketing or other business communication that's focused on selling what it's saying. A bullet-pointy style is used in many kinds of business communication. Etc.

It makes me wonder if part of what happened was a kind of melding of common styles from several different kinds of writing.

Re: Why does Opus 5 feel worse to work with?

#894
post #428

Earlier quoted context omitted.

Everything that claude writes fits into the same aesthetic structure. The aesthetic is that of an expert slowly revealing an insight to the user. The actual content doesn't matter. - "Introduction that rephrases your prompt." - "3 paragraphs, with one section of bullet points" - "The Twist" - "The Bottom Line" It's really obvious once you see it. Every single prompt, from a quantum physics question to a mundane obser…

You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).

- load-bearing

- smoking gun

- I found the seam

Re: Why does Opus 5 feel worse to work with?

#895

A gem Opus 5 gifted to me today: "A devastating pair of findings, and the first is beautiful in a way worth naming: the anti-vacuity floor is what blinds the gate to a vacuous case."

This thread seems like the wort of folk who might get a laugh at https://clanker-quotes.com/

> Good instinct - let me pull the threads together, because the loose pair and the encode_frame overlap turn out to be the same problem wearing two hats.

This one is one of my all time favorites. There’s just something about a code problem “wearing two hats” that absolutely kills me. I crack up every time…

Re: Why does Opus 5 feel worse to work with?

#896

Earlier quoted context omitted.

I didn't make any claim about being smart or not aside from brief mockery of people who care about the word and things like IQ tests I'd personally trust a rigorous, well-evinced, and methodical dissection of a topic by someone with completely average intelligence over a lazy broad generalization by someone who scored high on a paper test that people (wrongly) assume maps onto the ability to assess and analyze the wo…

I've participated in peer review from both sides, and it's much less impressive than people imagine from the outside.

[dead]

Re: Why does Opus 5 feel worse to work with?

#897

Earlier quoted context omitted.

see sibling comment, > see it actually improve just through accreting context this actually happens and has been tested.

> see sibling comment, > > see it actually improve just through accreting context > this actually happens and has been tested. I specifically said a novel task outside of the explicit training. And I already agreed that the so-called thinking models do some level of logical reasoning. But being able to engage in some level of reasoning because it has learned logical inference rules doesn't mean it's actually thinking…

> outside of explicit training

what does this mean anymore? i can invent a proof assistant language that was not in the training set and the llm does a fantastic job with it.

(github.com/ityonemo/bpa)

well it learned some sort of programming language and some sort of logic

well can you not see that those sorts of analogies can effectively make nothing in the knowable universe out of distribution? if you did such a categorization for human learning you could likewise say, "humans never can do anything outside of training set" too.

Re: Why does Opus 5 feel worse to work with?

#898

Earlier quoted context omitted.

My personal bugbear is its usage of "grain" where normally you'd use "granularity", if at all.

Lately I've noticed everything is "sharp" or it "sharpens" the point or question or something.

That's a sharp question, and it reframes the whole point.

Re: Why does Opus 5 feel worse to work with?

#899

Earlier quoted context omitted.

You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).

I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…

It's endlessly annoying to me that the em-dash has become the canary in the coalmine of AI writing because I've always used them extremely liberally in my writing.

Honestly, this might sound elitist, but I suspect it's because it's an "advanced" punctuation that is not known by most people, so is not commonly used. But the training corpus of these models puts more weight of academic writing or published books writing where the em-dash is much more commonly used.

Re: Why does Opus 5 feel worse to work with?

#900

Earlier quoted context omitted.

You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).

I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…

I think it presses a few buttons we probably recognize (if subconsciously) and find distasteful.

The verbosity makes me think of two things in particular:

- The classic essay written by someone who has 125 words worth of actual content but a 1500 word minimum. Those three paragraphs could be bullet points and convey the meaning just as well. The screen-filling chart of every test case you ran that came back green manages to be less actionable than a direct "one test out of 54 failed." I fully expect to see Claude tell us that "Support Ticket 8257 is a Land Of Contrasts" at some point.

- The sitcom trope of the person caught in a lie who figures if they can keep adding more and more detail he'll be believed and can escape the awkward conversation. Stop. Just stop. You're proposing a fix on a CODEBASE THE CUSTOMER DOES NOT EVEN USE. Cue laugh track, cut to commercial.

Post reply on HN