Benchmarking Opus 5 on SlopCodeBench
81–90 of 131 posts
Re: Benchmarking Opus 5 on SlopCodeBench
#82The original paper cited by this post does try to see if improved prompting will make a big difference in the end using a `plan_first` prompt variant, but find no influence on pass rate at the end of the benchmark. The `plan_first` seems to assume coding agents will just refactor once they finish features, but I don't think they tend to refactor significantly unless they are told to fix bugs rather than build features. The benchmark leaves tests hidden with no fail-to-pass feedback, so that may be why degradation is monotonic.
Re: Benchmarking Opus 5 on SlopCodeBench
#83Really hoping that all of the attention you're bringing to the longitudinal sloppification of codebases makes it back to the labs and creates some pressure to improve that trait of the models. This new benchmark seems promising. At the same time, I imagine it will be hard for them to prioritize this over improving flashy one-shots of impressive zero-to-one feats that demo so well and attract more customers. Can we ha…
Re: Benchmarking Opus 5 on SlopCodeBench
#84I always have Claude recite a pledge before starting coding to fix redundant code it notices over time. It does seem to find redundancies, but only when I point out bugs, that's when it goes into fixing mode and actually applies my Don't Repeat Yourself preference from the CLAUDE.md. The original paper cited by this post does try to see if improved prompting will make a big difference in the end using a `plan_first`…
Re: Benchmarking Opus 5 on SlopCodeBench
#85Re: Benchmarking Opus 5 on SlopCodeBench
#86I always have Claude recite a pledge before starting coding to fix redundant code it notices over time. It does seem to find redundancies, but only when I point out bugs, that's when it goes into fixing mode and actually applies my Don't Repeat Yourself preference from the CLAUDE.md. The original paper cited by this post does try to see if improved prompting will make a big difference in the end using a `plan_first`…
This is just superstition.
Re: Benchmarking Opus 5 on SlopCodeBench
#87This matches my experience of Opus 5 being a nice improvement over Opus 4.8, but not being revolutionary like Fable felt. I’ve now replaced my use of Opus 4.8 xhigh with Opus 5 medium, and I’m using less tokens and it’s quicker. I can understand people being annoyed by its writing style but for getting work done that really doesn’t bother me. I’ve been really enjoying using it.
Can you elaborate on what felt revolutionary to you about Fable?
Re: Benchmarking Opus 5 on SlopCodeBench
#88If the code is ugly, but defects are low—does it matter?
If the code is hard to read, but clients are happy—should I care?
Genuinely asking. Weird times we live in.
Re: Benchmarking Opus 5 on SlopCodeBench
#89I may be crucified for asking this but: is there any proof that slop matters beyond our sensibilities as developers? If the code is ugly, but defects are low—does it matter? If the code is hard to read, but clients are happy—should I care? Genuinely asking. Weird times we live in.
Re: Benchmarking Opus 5 on SlopCodeBench
#90Earlier quoted context omitted.
This is just superstition.
“Coding” can be more like communing with an otherworldly presence via esoteric gesturing nowadays than ever. Superstition leads countries and companies.