Earlier quoted context omitted.
[flagged]
How is this useful or insightful? You ever go to forums full of entomology specialists and tell them you don’t understand their fancy terms?
Benchmarking Opus 5 on SlopCodeBench
11–20 of 131 posts
Re: Benchmarking Opus 5 on SlopCodeBench
#12Re: Benchmarking Opus 5 on SlopCodeBench
#13Opus 5 is an overconfident stupid model. It tries to generate too much slop, tries to act like everything will fall. I have reversed back to fable and codex sol.
I did find opus 5 quite handy for general knowledge work and visual design, without the cost of fable (e.g. the graphics in this post are made by opus 5)
but its not noticeably better than opus 4.8 in those regards, and I would not miss it if forced to go back to 4.8
Re: Benchmarking Opus 5 on SlopCodeBench
#14Re: Benchmarking Opus 5 on SlopCodeBench
#15Earlier quoted context omitted.
How is this useful or insightful? You ever go to forums full of entomology specialists and tell them you don’t understand their fancy terms?
My point is that the differences between these models are so minor that obsessively benchmarking them comes across as navel-gazing.
Re: Benchmarking Opus 5 on SlopCodeBench
#16Earlier quoted context omitted.
How is this useful or insightful? You ever go to forums full of entomology specialists and tell them you don’t understand their fancy terms?
My point is that the differences between these models are so minor that obsessively benchmarking them comes across as navel-gazing.
If a model isn’t a step function change? Welcome to research.
Re: Benchmarking Opus 5 on SlopCodeBench
#17It's especially relevant now that models are good enough to solve ~most point-in-time problems.
Some relevant but disconnected thoughts:
- deterministic scores are so nice
- what "maintainable" is is probably some high dimensional space described by these signals; it'd probably require some human labeling to figure out where this space is
- another signal I've been thinking about and I'm seeing increasingly get brought up is the state space of a system; I'm seeing formal methods pop up a lot recently
Re: Benchmarking Opus 5 on SlopCodeBench
#18I’ve now replaced my use of Opus 4.8 xhigh with Opus 5 medium, and I’m using less tokens and it’s quicker. I can understand people being annoyed by its writing style but for getting work done that really doesn’t bother me. I’ve been really enjoying using it.
Re: Benchmarking Opus 5 on SlopCodeBench
#19Nice! I actually ran across this paper+benchmark recently, too. It's the first I've found that start to aim at some of the non-functional and longitudinal requirements that I think have always been an important part of writing production code. It's especially relevant now that models are good enough to solve ~most point-in-time problems. Some relevant but disconnected thoughts: - deterministic scores are so nice - wh…
this is a nicely succinct way to put this - a multi-dimensional space where no single metric is really useful
state space of the system is interesting too. I would guess that for any production software with dependencies like databases/third parties that might be too hard to measure, but if you can silo off parts of your system into bounded state machines, it may be a value metric on some module behind a clean interface.
I think the kubernetes control loop model is a great instance of this, a handful of scoped components that own a control loop across a well-defined state machine, that can operate / recover in the face of most network partitions or downtime - the promise of CRDTs but rather more a pragmatic approach to it
Re: Benchmarking Opus 5 on SlopCodeBench
#20Opus 5 is an overconfident stupid model. It tries to generate too much slop, tries to act like everything will fall. I have reversed back to fable and codex sol.
yes sol is still my daily driver for most coding tasks I did find opus 5 quite handy for general knowledge work and visual design, without the cost of fable (e.g. the graphics in this post are made by opus 5) but its not noticeably better than opus 4.8 in those regards, and I would not miss it if forced to go back to 4.8