Benchmarking Opus 5 on SlopCodeBench
github.com
Benchmarking Opus 5 on SlopCodeBench
1–10 of 131 posts
Re: Benchmarking Opus 5 on SlopCodeBench
#2Re: Benchmarking Opus 5 on SlopCodeBench
#3Re: Benchmarking Opus 5 on SlopCodeBench
#4I haven't joined your chats in a while but glad to see you put this together, I truly feel as though opus 5 is not much of an improvement. The only time i ever felt a wow factor was opus 4, 4.6 and fable pre trump admin lobotimizing
I'm hoping to do some more work with sol/fable in the mix as well as exploring more languages and curating the problem set to include more of the benchmark
I also kinda felt like opus4.5 was dumber than 4.1 personally, maybe a little biased since 4.5 was 2.5x faster and 2.5x cheaper seems to indicate its a smaller model
Re: Benchmarking Opus 5 on SlopCodeBench
#5finally the benchmark for me
Re: Benchmarking Opus 5 on SlopCodeBench
#6I haven't joined your chats in a while but glad to see you put this together, I truly feel as though opus 5 is not much of an improvement. The only time i ever felt a wow factor was opus 4, 4.6 and fable pre trump admin lobotimizing
yeah this was just a start - the fastest cheapest thing we could try for a brand new model. I'm hoping to do some more work with sol/fable in the mix as well as exploring more languages and curating the problem set to include more of the benchmark I also kinda felt like opus4.5 was dumber than 4.1 personally, maybe a little biased since 4.5 was 2.5x faster and 2.5x cheaper seems to indicate its a smaller model
Re: Benchmarking Opus 5 on SlopCodeBench
#7Re: Benchmarking Opus 5 on SlopCodeBench
#8I haven't joined your chats in a while but glad to see you put this together, I truly feel as though opus 5 is not much of an improvement. The only time i ever felt a wow factor was opus 4, 4.6 and fable pre trump admin lobotimizing
Re: Benchmarking Opus 5 on SlopCodeBench
#9Did you not benchmark latest GPT 5.6 or GLM 5.1/Kimi K3 because of cost? I can run them if you share how you ran them
> fetch this article for slopcodebench and help me run an eval on a subset of problems with opus 5 https://arxiv.org/html/2603.24755v1 > Get all the context, fetch any mentioned repos, and then propose a plan to me.
> i have an anthropic API key in .... > Let's do the three challenges with Opus 4.8 and Opus 5 and Fable please. I like your minimal set. Let's try it. What do you need from me?
> Actually I changed my mind. I want to do two of the easy ones you picked and then I want you to pick the one with more checkpoints, maybe one of the harder ones, not the very hardest one but one with a higher number of checkpoints.
> Actually let's do one easy, one medium, and one hard problem please. If we have a hard problem I'd like to see that.
> lets rock - i think lets just do opus 4.8 and sonnet 5 and opus 5 since we our ZDR will block fable
Re: Benchmarking Opus 5 on SlopCodeBench
#10I haven't joined your chats in a while but glad to see you put this together, I truly feel as though opus 5 is not much of an improvement. The only time i ever felt a wow factor was opus 4, 4.6 and fable pre trump admin lobotimizing
[flagged]
You ever go to forums full of entomology specialists and tell them you don’t understand their fancy terms?