Live data from Hacker News

Why does Opus 5 feel worse to work with?

mun-logadan.github.io

401–410 of 915 posts

Re: Why does Opus 5 feel worse to work with?

#401

I've gone back to 4.8. 5 would constantly veer of in random directions if not working from 100% strict and narrow instructions. I find it weird there's not more discussion here on HN on how the most used model now has clearly degraded in quality and it seems we've hit a peak and are on a downslope - because the model is clearly smaller or more economical for Anthropic no doubt about it, and the benchmaxxing they do i…

> I find it weird there's not more discussion here on HN on how the most used model now has clearly degraded in quality and it seems we've hit a peak and are on a downslope It's not weird, because it's an anecdote, not an accepted fact. Personally I've not been too happy with Opus 5, but I've had similar experiences with other models previously, feeling like they didn't quite fit with my working style. So nothing ind…

> nothing indicates we've hit a peak

Opus 5 is arguably a regression but GPT 5.6 is pretty strong evidence that we haven't hit a peak. I think I actually prefer Sol to Fable at this point.

Re: Why does Opus 5 feel worse to work with?

#402
post #353

Earlier quoted context omitted.

Anthropic is lucky that they've built a lot of loyalty over the last year that they can burn through right now. I see people talking about switching back to Opus 4.8 rather that using 5.6 Sol, which is wild. My current approach is to occasionally use Fable for high-intelligence tasks but use Sol as the translator and clean-upper afterwards, and otherwise just use Sol for everything. Fable sometimes says the most insa…

I'm still on opus 4.6 for a healthy chunk on work; the technical competence has lagged behind but the slop-comment generation and misdirected self-initiated actions on newer models ultimately burn more time than a little more babysitting, but I'm optimizing for minimized slop generation over sheer generation speed.

Yeah I'm using 5 for its technical abilities, but I much preferred 4.6's personality. 5 loved to double check everything, including the double-checks, and I have to stop it and tell it "this is irrelevant" or "this is out of scope" all the time, or sometimes I'll go leave it to do a task and come back and it's still verifying the tiniest details of its assumptions before actually doing anything

Re: Why does Opus 5 feel worse to work with?

#403

My latest trick (literally from yesterday) is to just ask it to write according to ISO 24495-1, the standard for plain language: > [This standard is] for anybody who creates or helps create documents. The widest use of plain language is for documents that are intended for the general public. However, it is also applicable, for example, to technical writing, legislative drafting or using controlled languages. You don'…

Keep in mind all this kind of stuff can make the model less capable. If it has to think in "plain" English, it may well be squashing quality of code etc output. I'm not sure how true this is, but when using "forced" json output it def had a big drop off in quality - https://arxiv.org/html/2408.02442v3 . I think you're better not fighting it with hacks like this and find a different model.

I would not overgeneralize from paper. Firstly: forcing JSON output is, in my opinion, a bigger change than asking it to match the above style guidelines, and secondly, as is always the case with these kinds of papers, what was true for the model tested in the paper may either be completely false, or greatly reduced, in later models. That paper is almost 2 years old and models today have been trained in very different ways (or more accurately post trained in very different ways) and are in general far more capable.

Based on that paper, I would maybe try to check if it was true for a modern use case, I would very much not assume it was still true.

Re: Why does Opus 5 feel worse to work with?

#404

Earlier quoted context omitted.

the tip that was floating around on x was to tell it to use "ASD-STE100 Simplified Technical English" cladue desktop has an instructions sections under general options, you can put something like "try to stick to ASD-STE100 Simplified Technical English, keep answers short and to the point" funnily enough the placeholder they suggest when its empty is "keep answers short and to the point"

CLAUDE.md is mostly powerless against the reinforcement learned crap. I'm up to three separate instructions telling it to cut out the hyper verbose, retelling history comments and it still writes them every time.

Try spacing them out instead. I.e. a mini-workflow with a self-review step. Works for both planning and coding.

Re: Why does Opus 5 feel worse to work with?

#405
post #28

Earlier quoted context omitted.

From the little i understand that wouldnt be an issue because the model is ‘just’ using interchangeable words in a mathematical non-random way. Like using the same number of adjectives and the exct same words, but in a order that wouldn’t be mathematically plausible unless it was the watermark

I wouldn’t exactly put it like that. It’s moreso the model sometimes outputting non-optimal tokens in a way that’s detectable if you know the algorithm. It seems possible for that to make the response “drift” far from what it would’ve been, because it’s constant entropy that adds up after time. (However, according to Anthropic and Google, it doesn’t really impact the quality of responses. I find that a bit hard to be…

> according to Anthropic and Google, it doesn’t really impact the quality of responses. I find that a bit hard to believe, although those guys are much smarter than I.

I would guess "it doesn't impact the quality of responses" was guaranteed to be claimed before they even implemented any of the watermarking.

And would come from marketing, not the people who implemented it.

Re: Why does Opus 5 feel worse to work with?

#407
post #406

Is it possible that the same version performs worse after time? I used Opus 4.8 two to three months ago for writing a paper and I swear the responses and the output was MUCH better than in July.

Maybe they want you to feel how much better Opus 5 is :)

Re: Why does Opus 5 feel worse to work with?

#408
Yes, definitely. It bounces around, goes off and does its own thing, etc. From Opus 4.8, it seems to have blindly increased its confidence while reducing its focus and efficiency. I found it so hard to corral that I reverted back to Opus 4.8. I was constantly having to refocus and redirect Opus 5. It was like an eager intern.

Re: Why does Opus 5 feel worse to work with?

#409
post #377
post #353

Earlier quoted context omitted.

Anthropic is lucky that they've built a lot of loyalty over the last year that they can burn through right now. I see people talking about switching back to Opus 4.8 rather that using 5.6 Sol, which is wild. My current approach is to occasionally use Fable for high-intelligence tasks but use Sol as the translator and clean-upper afterwards, and otherwise just use Sol for everything. Fable sometimes says the most insa…

I think Fable is the beginning of Anthropic switching to training models as agent-first, tool second. It’s certainly the best model if you want something to work autonomously without supervision and don’t care to read the code. The code and writing is ugly but it can complete huge tasks and fix its own work.

I thought having a model 'fix its own work' leads to model collapse?

Re: Why does Opus 5 feel worse to work with?

#410

Anthropic, if you're listening - by the time this crops up on Reddit, the front page of HN, etc.... you should be expecting calls from CEOs of major corporations next threatening to abandon ship... We've seen this pattern before several times.. I hope they are listening and address this publicly. I'm not sure what is going on, some users report it works fine or great, others report the degradation. I've experienced b…

You are expecting consistent QoS from a randomly sampled mathematical function.

No, they're expecting to see a failure rate consistent with previous failure rates, not periods of low failure rates and other periods of high failure rates, with the same model.

And you can expect more consistency from SOTA models than you can from an old model like GPT3--you agree, right?

GP expects the same level of consistency throughout their time using the same model. Not request-to-request, more like day-to-day.

Post reply on HN