Live data from Hacker News

StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

app.uniclaw.ai

21–30 of 97 posts

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#22
post #5

According to openrouter.ai it looks like StepFun 3.5 Flash is the most popular model at 3.5T tokens, vs GLM 5 Turbo at 2.5T tokens. Claude Sonnet is in 5th place with 1.05T tokens. Which isn't super suprising as StepFun is ~about 5% the price of Sonnet. https://openrouter.ai/apps?url=https%3A%2F%2Fopenclaw.ai%2F

> the most popular model It was free for a long time. That usually skews the statistics. It was the same with grok-code-fast1.

Exactly. When I read the headline I thought: "Ofc it is, its free."

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#23

Earlier quoted context omitted.

all 300+ battle data are available at https://app.uniclaw.ai/arena/battles , every single battle is shown with raw conversional history, produced files, judge's verdict and final scores

Thanks! Is the judge an LLM? There's lot of references to "just like LMArena", but LMArena is human evaluated?

> Is the judge an LLM?

Yes, judge is one of opus 4.6, gpt 5.4, gemini 3.1 pro (submitter can choose). Self judge (judge model is also one of the participants) is excluded when computing ranking.

> There's lot of references to "just like LMArena", but LMArena is human evaluated?

Yeah LMArena is human evaluated, but here i found it not practical to gather enough human evaluation data because the effort it take to compare the result is much higher:

- for code, judge needs to read through it to check code quality, and actually run it to see the output

- when producing a webpage or a document, judge needs to check the content and layout visually

- when anything goes wrong, judge needs to read the execution log to see whether partial credit shall be granted

if you look at the cost details of each battle (available at the bottom of battle detail page), judge typically cost more than any participant model.

if we evaluate with human, i would say each evaluation can easily take ~5-10 min

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#25

Earlier quoted context omitted.

Thanks! Is the judge an LLM? There's lot of references to "just like LMArena", but LMArena is human evaluated?

> Is the judge an LLM? Yes, judge is one of opus 4.6, gpt 5.4, gemini 3.1 pro (submitter can choose). Self judge (judge model is also one of the participants) is excluded when computing ranking. > There's lot of references to "just like LMArena", but LMArena is human evaluated? Yeah LMArena is human evaluated, but here i found it not practical to gather enough human evaluation data because the effort it take to compa…

Fair enough, yeah, agent evals are hard especially across N models :/

Thanks for replying btw, didn't mean any disrespect, good on you for not getting aggro about feedback

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#26

I ran 300+ benchmarks across 15 models in OpenClaw and published two separate leaderboards: performance and cost-effectiveness. The two boards look nothing alike. Top 3 performance: Claude Opus 4.6, GPT-5.4, Claude Sonnet 4.6. Top 3 cost-effectiveness: StepFun 3.5 Flash, Grok 4.1 Fast, MiniMax M2.7. The most dramatic split: Claude Opus 4.6 is #1 on performance but #14 on cost-effectiveness. StepFun 3.5 Flash is #1 co…

Could you add a column for time or number of tokens? Some models take forever because of their excessive reasoning chains.

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#27
post #22

Earlier quoted context omitted.

> the most popular model It was free for a long time. That usually skews the statistics. It was the same with grok-code-fast1.

Exactly. When I read the headline I thought: "Ofc it is, its free."

I should have clarified I didn't use the free version...

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#28

I ran 300+ benchmarks across 15 models in OpenClaw and published two separate leaderboards: performance and cost-effectiveness. The two boards look nothing alike. Top 3 performance: Claude Opus 4.6, GPT-5.4, Claude Sonnet 4.6. Top 3 cost-effectiveness: StepFun 3.5 Flash, Grok 4.1 Fast, MiniMax M2.7. The most dramatic split: Claude Opus 4.6 is #1 on performance but #14 on cost-effectiveness. StepFun 3.5 Flash is #1 co…

>Other surprises: GLM-5 Turbo, Xiaomi MiMo v2 Pro, and MiniMax M2.7 all outrank Gemini 3.1 Pro on performance

This has also been my subjective experience But has also been objective in terms of cost.

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#29

I ran 300+ benchmarks across 15 models in OpenClaw and published two separate leaderboards: performance and cost-effectiveness. The two boards look nothing alike. Top 3 performance: Claude Opus 4.6, GPT-5.4, Claude Sonnet 4.6. Top 3 cost-effectiveness: StepFun 3.5 Flash, Grok 4.1 Fast, MiniMax M2.7. The most dramatic split: Claude Opus 4.6 is #1 on performance but #14 on cost-effectiveness. StepFun 3.5 Flash is #1 co…

Could you add a column for time or number of tokens? Some models take forever because of their excessive reasoning chains.

both are shown in battle detail page already. Time is shown in Scores table. Number of tokens are shown in Cost details at the bottom of the Scores. (I thought most people just want to see cost in USD so I put token details at the bottom)

Re: StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

#30

Earlier quoted context omitted.

> Is the judge an LLM? Yes, judge is one of opus 4.6, gpt 5.4, gemini 3.1 pro (submitter can choose). Self judge (judge model is also one of the participants) is excluded when computing ranking. > There's lot of references to "just like LMArena", but LMArena is human evaluated? Yeah LMArena is human evaluated, but here i found it not practical to gather enough human evaluation data because the effort it take to compa…

Fair enough, yeah, agent evals are hard especially across N models :/ Thanks for replying btw, didn't mean any disrespect, good on you for not getting aggro about feedback

I appreciate honest feedback, best way to learn :)
Post reply on HN