Live data from Hacker News

Benchmarking Opus 5 on SlopCodeBench

github.com

1–10 of 131 posts

Re: Benchmarking Opus 5 on SlopCodeBench

#4

I haven't joined your chats in a while but glad to see you put this together, I truly feel as though opus 5 is not much of an improvement. The only time i ever felt a wow factor was opus 4, 4.6 and fable pre trump admin lobotimizing

yeah this was just a start - the fastest cheapest thing we could try for a brand new model.

I'm hoping to do some more work with sol/fable in the mix as well as exploring more languages and curating the problem set to include more of the benchmark

I also kinda felt like opus4.5 was dumber than 4.1 personally, maybe a little biased since 4.5 was 2.5x faster and 2.5x cheaper seems to indicate its a smaller model

Re: Benchmarking Opus 5 on SlopCodeBench

#6
post #4

I haven't joined your chats in a while but glad to see you put this together, I truly feel as though opus 5 is not much of an improvement. The only time i ever felt a wow factor was opus 4, 4.6 and fable pre trump admin lobotimizing

yeah this was just a start - the fastest cheapest thing we could try for a brand new model. I'm hoping to do some more work with sol/fable in the mix as well as exploring more languages and curating the problem set to include more of the benchmark I also kinda felt like opus4.5 was dumber than 4.1 personally, maybe a little biased since 4.5 was 2.5x faster and 2.5x cheaper seems to indicate its a smaller model

yeah i agree

Re: Benchmarking Opus 5 on SlopCodeBench

#9

Did you not benchmark latest GPT 5.6 or GLM 5.1/Kimi K3 because of cost? I can run them if you share how you ran them

no i'm spinning those up at some point this week. here's the first few prompts I used (claude opus 5 as the research orchestrator), (these were interspersed with lots of tools and assistant messages but it should get you kicked off.

> fetch this article for slopcodebench and help me run an eval on a subset of problems with opus 5 https://arxiv.org/html/2603.24755v1 > Get all the context, fetch any mentioned repos, and then propose a plan to me.

> i have an anthropic API key in .... > Let's do the three challenges with Opus 4.8 and Opus 5 and Fable please. I like your minimal set. Let's try it. What do you need from me?

> Actually I changed my mind. I want to do two of the easy ones you picked and then I want you to pick the one with more checkpoints, maybe one of the harder ones, not the very hardest one but one with a higher number of checkpoints.

> Actually let's do one easy, one medium, and one hard problem please. If we have a hard problem I'd like to see that.

> lets rock - i think lets just do opus 4.8 and sonnet 5 and opus 5 since we our ZDR will block fable

Re: Benchmarking Opus 5 on SlopCodeBench

#10
post #8

I haven't joined your chats in a while but glad to see you put this together, I truly feel as though opus 5 is not much of an improvement. The only time i ever felt a wow factor was opus 4, 4.6 and fable pre trump admin lobotimizing

[flagged]

How is this useful or insightful?

You ever go to forums full of entomology specialists and tell them you don’t understand their fancy terms?

Post reply on HN