Live data from Hacker News

StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

app.uniclaw.ai

61–70 of 97 posts

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#61
post #21

why do half the comments here read like ai trying to boost some sort of scam?

Because there's absolutely nothing stopping that from happening. There are bots on Reddit, there are of course bots on here, a VPN friendly site where you don't even need an email. But a lot of people don't want to admit it.

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#62
post #50

I was excited to read through this to find out how these tasks are evaluated at scale. Lots of scary looking formulas with sigmas and other Greek letters. Then I clicked on one task to see what it looks like “on the ground”: https://app.uniclaw.ai/arena/DDquysCGBsHa (not cherry picked- literally the first one I clicked on) The task was: > Find rental properties with 10 bedrooms and 8 or more bathrooms within a 1 hour…

"commiserate" - did you mean "commensurate"?

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#63

Earlier quoted context omitted.

both are shown in battle detail page already. Time is shown in Scores table. Number of tokens are shown in Cost details at the bottom of the Scores. (I thought most people just want to see cost in USD so I put token details at the bottom)

I would have liked aggregated results instead. Expanding 300 tables is a bit tiresome. But I guess that is easy with AI now. Here is a scatter plot of quality vs duration https://i.imgur.com/wFVSpS5.png and quality vs cost https://i.imgur.com/fqM4edw.png But I just noticed that my plot is meaningless because it conflates model quality with provider uptime. Claude Haiku has a higher average quality than Claude Opus, w…

i added native plot and stats for aggregated results, on arena page. please check it out!

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#64

Earlier quoted context omitted.

both are shown in battle detail page already. Time is shown in Scores table. Number of tokens are shown in Cost details at the bottom of the Scores. (I thought most people just want to see cost in USD so I put token details at the bottom)

I would have liked aggregated results instead. Expanding 300 tables is a bit tiresome. But I guess that is easy with AI now. Here is a scatter plot of quality vs duration https://i.imgur.com/wFVSpS5.png and quality vs cost https://i.imgur.com/fqM4edw.png But I just noticed that my plot is meaningless because it conflates model quality with provider uptime. Claude Haiku has a higher average quality than Claude Opus, w…

The second chart depicts StepFun > Sonnet > Opus in quality?

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#65
post #50

I was excited to read through this to find out how these tasks are evaluated at scale. Lots of scary looking formulas with sigmas and other Greek letters. Then I clicked on one task to see what it looks like “on the ground”: https://app.uniclaw.ai/arena/DDquysCGBsHa (not cherry picked- literally the first one I clicked on) The task was: > Find rental properties with 10 bedrooms and 8 or more bathrooms within a 1 hour…

"commiserate" - did you mean "commensurate"?

At that point commiserations were in order

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#66
post #50

I was excited to read through this to find out how these tasks are evaluated at scale. Lots of scary looking formulas with sigmas and other Greek letters. Then I clicked on one task to see what it looks like “on the ground”: https://app.uniclaw.ai/arena/DDquysCGBsHa (not cherry picked- literally the first one I clicked on) The task was: > Find rental properties with 10 bedrooms and 8 or more bathrooms within a 1 hour…

"commiserate" - did you mean "commensurate"?

Sorry, yes. I was typing quickly

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#67
post #64

Earlier quoted context omitted.

I would have liked aggregated results instead. Expanding 300 tables is a bit tiresome. But I guess that is easy with AI now. Here is a scatter plot of quality vs duration https://i.imgur.com/wFVSpS5.png and quality vs cost https://i.imgur.com/fqM4edw.png But I just noticed that my plot is meaningless because it conflates model quality with provider uptime. Claude Haiku has a higher average quality than Claude Opus, w…

The second chart depicts StepFun > Sonnet > Opus in quality?

check out my reply, his chart is plotting the wrong metric (average quality score)

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#68

None of the Qwen 3.5 models seem present? I’ve heard people are pretty happy with the smaller 3.5 versions. I would be curious to see those too. I would also be interested to see "KAT-Coder-Pro-V2" as they brag about their benchmarks in these bots as well

If they use OpenRouter pricing then the Qwen3.5 models are going to be poor value.

The Qwen3.5 27B model on OR is $1.56/million tokens out (it used to be $2.4/mil).

Meanwhile Minimax M2.7 (a much larger model) is $1.2/mil out.

The smaller and medium tier Qwen3.5 models are only really cost effective if you run them yourself.

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#69

another thing from the bench I didn't expect: gemini 3.1 pro is very unreliable at using skills. sometimes it just reads the skill and decide to do nothing, while opus/sonnet 4.6 and gpt 5.4 never have this issue.

[flagged]

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#70
post #68

None of the Qwen 3.5 models seem present? I’ve heard people are pretty happy with the smaller 3.5 versions. I would be curious to see those too. I would also be interested to see "KAT-Coder-Pro-V2" as they brag about their benchmarks in these bots as well

If they use OpenRouter pricing then the Qwen3.5 models are going to be poor value. The Qwen3.5 27B model on OR is $1.56/million tokens out (it used to be $2.4/mil). Meanwhile Minimax M2.7 (a much larger model) is $1.2/mil out. The smaller and medium tier Qwen3.5 models are only really cost effective if you run them yourself.

Is Minimax M2.7 better than Qwen3.5 27B, or is it just bigger?
Post reply on HN