Benchmarking Opus 5 on SlopCodeBench
31–40 of 131 posts
Re: Benchmarking Opus 5 on SlopCodeBench
#32Earlier quoted context omitted.
How is this useful or insightful? You ever go to forums full of entomology specialists and tell them you don’t understand their fancy terms?
My point is that the differences between these models are so minor that obsessively benchmarking them comes across as navel-gazing.
Re: Benchmarking Opus 5 on SlopCodeBench
#33At the same time, I imagine it will be hard for them to prioritize this over improving flashy one-shots of impressive zero-to-one feats that demo so well and attract more customers.
Can we have simonw make "pelican on a bicycle after 1000 requests for iteration" popular?
Re: Benchmarking Opus 5 on SlopCodeBench
#34Opus 5 is an overconfident stupid model. It tries to generate too much slop, tries to act like everything will fall. I have reversed back to fable and codex sol.
I'm not ready to blame Opus 5 for being stupid. Perhaps we have a prompt buried somewhere that's essentially asking it to be pedantic, and it's just obeying the prompt.
Re: Benchmarking Opus 5 on SlopCodeBench
#35Really hoping that all of the attention you're bringing to the longitudinal sloppification of codebases makes it back to the labs and creates some pressure to improve that trait of the models. This new benchmark seems promising. At the same time, I imagine it will be hard for them to prioritize this over improving flashy one-shots of impressive zero-to-one feats that demo so well and attract more customers. Can we ha…
yes the labs will always prioritize the vibeslop dopamine casino as far as I can tell - making the models useful and addictive for unsophisticated users, sometimes at the expense or at the very least at the ignorance of the needs of power users
Re: Benchmarking Opus 5 on SlopCodeBench
#36This matches my experience of Opus 5 being a nice improvement over Opus 4.8, but not being revolutionary like Fable felt. I’ve now replaced my use of Opus 4.8 xhigh with Opus 5 medium, and I’m using less tokens and it’s quicker. I can understand people being annoyed by its writing style but for getting work done that really doesn’t bother me. I’ve been really enjoying using it.
Medium vs High? Why? From all the charts I've seen the performance jump is pretty large from med -> high (not as noticeable from high -> xhigh).
Quality vs cost - medium is the sweet (perhaps better too!) spot.
Re: Benchmarking Opus 5 on SlopCodeBench
#37Earlier quoted context omitted.
Noticed that too. I wonder if these things just degrade over time, perhaps with the way it writes memories about my project as it goes
I’ve observed the degradation, but I suspect what’s happening is they’re tuning it for lower inference costs. Maybe turning down the amount of thinking, maybe quantizing, maybe something else. It seems like there’s a week by week and sometimes day by day change in performance when on a subscription plan using their harnesses.
Re: Benchmarking Opus 5 on SlopCodeBench
#38Earlier quoted context omitted.
Medium vs High? Why? From all the charts I've seen the performance jump is pretty large from med -> high (not as noticeable from high -> xhigh).
https://cognition.com/frontiercode Quality vs cost - medium is the sweet (perhaps better too!) spot.
Re: Benchmarking Opus 5 on SlopCodeBench
#39To what degree is this a harness/system prompt problem? Models maybe should implement new stuff with as little impact on the existing stuff as possible by default? A simple system prompt for it to always check the code after task completion for proper simplifications, abstractions and cleanups before returning to the user? Instructions to retain "story like" readability of the code.