Live data from Hacker News

Why does Opus 5 feel worse to work with?

mun-logadan.github.io

581–590 of 915 posts

Re: Why does Opus 5 feel worse to work with?

#581

It's not even code for me, but the prose it writes. For some reason, the way Opus 5 "talk" elicits frustration in a way that 4.5 to 4.8 never did. Can't put my finger on why, but I've flipped over to Codex because what it produced wasn't worth the frustration.

>Can't put my finger on why, but I've flipped over to Codex because what it produced wasn't worth the frustration.

Its because its hard to understand what it means and is outright incoherent at times. It has its own style that I cant describe well either but the bottom line is its hard to understand what the fuck its even trying to say. Reading nonsense is tyring.

Re: Why does Opus 5 feel worse to work with?

#582

Earlier quoted context omitted.

I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…

> I am curious why LLM writing has such an uncanny valley feel to it. Because they are HEAVILY trained to give addictive responses. They don't want to just answer your question. They want to sycophantically make you feel like a genius for being smart enough to use them.

This seems to be the load-bearing point that matters.

Re: Why does Opus 5 feel worse to work with?

#583

Earlier quoted context omitted.

> t doesn't think in humans the exact same behaviour (cheating) is slmost always the result of a chain of complex series of choices and environment-driven rationalization. if the llm doesn't cheat, you say "its just producing the most straightforward answer -- not thinking'. if it cheats, you say "weaseling out of hard thinking". damned if it cheats, damned if it doesn't. what evidence would convunce you that it is t…

> in humans the exact same behaviour (cheating) is slmost always the result of a chain of complex series of choices and environment-driven rationalization. It has been taught on the outcome of this. Broadly speaking, humans are lazy creatures (and when used judiciously, laziness is a good thing). For example: the famous example of Carmack not using a hashmap somewhere early on in, I think it was, Quake 1 initializati…

[deleted]

Re: Why does Opus 5 feel worse to work with?

#584
post #504

Earlier quoted context omitted.

So true. The breaking point for me was when it was constantly saying to wrap up because we'd done enough work for the day, but we had barely even done anything. Or constantly estimating that the next steps would take X number of weeks, and then knock it out in a single prompt. I demanded in no uncertain terms to stop giving pointless bogus estimates or telling me to stop working, and it just wouldn't. I also had a pr…

This!! I like to work weird hours of the night and Opus consistently likes to "wrap up" and say "it's been a long night" or "it's late" and "we've made great progress" It's infuriating, just do the work!

That is very strange. I havent seen that. It sounds like something leaking from its system prompt or something that its not handling well. Anthropic trying to prevent it from running longer or something.

Re: Why does Opus 5 feel worse to work with?

#585

Earlier quoted context omitted.

You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).

I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…

One reason is that one person’s idiosyncracies are limited in scope, but LLM-produced text is now everywhere. Also, filler words and mannerisms in speech we’re quite good at filtering out, but in written text the stand out much more.

Re: Why does Opus 5 feel worse to work with?

#586

Earlier quoted context omitted.

You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA. (I'm becoming allergic to how these things write).

I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily. In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds…

It’s the repetitiveness of style, the attempt to make everything seem as impactful as possible, the use of short sentences (too much Hemingway in the training data?), and obvious patterns like “it’s not this, it’s that” and several others.

Real human writing doesn’t follow such strict rules. When the same small set of rules is applied over and over throughout a text, it becomes obviously strange and machine-like.

Re: Why does Opus 5 feel worse to work with?

#588

Earlier quoted context omitted.

I have a chunky bit of functionality in my hobby app using babylon.js to render 3D worlds using things like portals and LoD rendering to manage the visual load. I built it out with a combo of Fable and Opus 5. I too got fed up with the prose of Opus in particular, and tried going back. Unfortunately, the previous models were less able to hack it. The prose was better but progress was worse. It wasn't just conversatio…

> "isSolidWall(x)" it would write something like "weightyNotEphemeral(x)" Is this a literal example? That is wild.

It is not the literal example because I told it to rename the function, but the function had the form thisNotThat for a boolean predicate that had a far more conventional name.

Re: Why does Opus 5 feel worse to work with?

#589

Earlier quoted context omitted.

I think I understand what you are trying to convey but I fear you've made the same mistake again, this time with "soul" instead of "intelligence." I think what you are getting at is that they are deterministic automata. They are machines. We have introduced randomness to add variation but it is an artificial randomness that simply perturbs the path traversed. When we choose words it isn't because of a token distribut…

I agree, and yet it is reasonable to ask if these notions of "context" (sorry!) - namely feelings, external world, body, senses, etc - are somehow distinct from an LLM's notions of context. Today they certainly capture different things, but given the right representations, why couldn't these human notions also be captured as "context"? The idea of memory does not seem to resolve this: if you allow the machine to "com…

Until now, human knowledge and values have built on prior human knowledge and experience. If AI is able to develop without human influence, I believe its value system will necessarily diverge into something alien.

Re: Why does Opus 5 feel worse to work with?

#590
I asked Borris at an Anthropic event in SF this week why Opus 5 and Fable 5 seem to forget so much when I give it rapid fire tasks when I'm reviewing a UI or something. He told me to run in safe mode, which didn't help at all.

My tinfoil hat theory is Anthropic is trying to get their new models to take on higher-level longer-running tasks, which has a trade-off against rapid-fire tactical use of an LLM.

For these reasons, I've always found the 5 series models from Anthropic aren't great and use 4.8 for a lot of my work.

Post reply on HN