Live data from Hacker News

Benchmarking Opus 5 on SlopCodeBench

github.com

81–90 of 131 posts

Re: Benchmarking Opus 5 on SlopCodeBench

#82
I always have Claude recite a pledge before starting coding to fix redundant code it notices over time. It does seem to find redundancies, but only when I point out bugs, that's when it goes into fixing mode and actually applies my Don't Repeat Yourself preference from the CLAUDE.md.

The original paper cited by this post does try to see if improved prompting will make a big difference in the end using a `plan_first` prompt variant, but find no influence on pass rate at the end of the benchmark. The `plan_first` seems to assume coding agents will just refactor once they finish features, but I don't think they tend to refactor significantly unless they are told to fix bugs rather than build features. The benchmark leaves tests hidden with no fail-to-pass feedback, so that may be why degradation is monotonic.

Re: Benchmarking Opus 5 on SlopCodeBench

#83

Really hoping that all of the attention you're bringing to the longitudinal sloppification of codebases makes it back to the labs and creates some pressure to improve that trait of the models. This new benchmark seems promising. At the same time, I imagine it will be hard for them to prioritize this over improving flashy one-shots of impressive zero-to-one feats that demo so well and attract more customers. Can we ha…

I remember reading a blog post a year or two ago about prompting like a printer driver and asking for increasingly ridiculous changes like time travel or wormhole capability. I don't know if I hallucinated that post but I haven't been able to find it again and I REALLY wish I could

Re: Benchmarking Opus 5 on SlopCodeBench

#84

I always have Claude recite a pledge before starting coding to fix redundant code it notices over time. It does seem to find redundancies, but only when I point out bugs, that's when it goes into fixing mode and actually applies my Don't Repeat Yourself preference from the CLAUDE.md. The original paper cited by this post does try to see if improved prompting will make a big difference in the end using a `plan_first`…

This is just superstition.

Re: Benchmarking Opus 5 on SlopCodeBench

#86

I always have Claude recite a pledge before starting coding to fix redundant code it notices over time. It does seem to find redundancies, but only when I point out bugs, that's when it goes into fixing mode and actually applies my Don't Repeat Yourself preference from the CLAUDE.md. The original paper cited by this post does try to see if improved prompting will make a big difference in the end using a `plan_first`…

This is just superstition.

“Coding” can be more like communing with an otherworldly presence via esoteric gesturing nowadays than ever. Superstition leads countries and companies.

Re: Benchmarking Opus 5 on SlopCodeBench

#87

This matches my experience of Opus 5 being a nice improvement over Opus 4.8, but not being revolutionary like Fable felt. I’ve now replaced my use of Opus 4.8 xhigh with Opus 5 medium, and I’m using less tokens and it’s quicker. I can understand people being annoyed by its writing style but for getting work done that really doesn’t bother me. I’ve been really enjoying using it.

Can you elaborate on what felt revolutionary to you about Fable?

Not OP, but to me it initially felt extremely proactive and energetic, just powering through roadblocks with ingenuity and enthusiasm. After it came back I was constantly getting refusals and downgrades for things Opus had been doing. I’ve written it off for my use cases and getting by just fine with Opus 4.8 and now 5.

Re: Benchmarking Opus 5 on SlopCodeBench

#88
I may be crucified for asking this but: is there any proof that slop matters beyond our sensibilities as developers?

If the code is ugly, but defects are low—does it matter?

If the code is hard to read, but clients are happy—should I care?

Genuinely asking. Weird times we live in.

Re: Benchmarking Opus 5 on SlopCodeBench

#89

I may be crucified for asking this but: is there any proof that slop matters beyond our sensibilities as developers? If the code is ugly, but defects are low—does it matter? If the code is hard to read, but clients are happy—should I care? Genuinely asking. Weird times we live in.

It's the eternal question with any kind of tech debt -- is it worth a little more velocity now in return for medium-to-long term slowdown? And there's no general right answer.

Re: Benchmarking Opus 5 on SlopCodeBench

#90
post #86

Earlier quoted context omitted.

This is just superstition.

“Coding” can be more like communing with an otherworldly presence via esoteric gesturing nowadays than ever. Superstition leads countries and companies.

Every day we get closer to being mystics having to commune with orbs in order to create phenomena.
Post reply on HN