GLM-5.3 Artificial Analysis Benchmarks
41–50 of 64 posts
Re: GLM-5.3 Artificial Analysis Benchmarks
#42Earlier quoted context omitted.
It would make reading and comparing a bit easier if the data was sorted by a dimension.
Cost per task: Model Score Cost / Task Output Tokens / Task ------------------------------------------------------------------------- Muse Spark 1.2 (xhigh) 56.8 $0.40 30,430 Gemini 3.7 Flash (high) 56.0 $0.40 36,847 GPT-5.6 Terra (max) 56.6 $0.51 20,838 GPT-5.6 Sol (high) 57.3 $0.52 7,545 GLM-5.2 (max) 53.0 $0.56 32,200 GLM-5.3 (max) 59.5 $0.68 41,107 GPT-5.5 (xhigh) 56.3 $0.69 16,893 Grok 4.6 (high) 60.9 $0.84 21,7…
Re: GLM-5.3 Artificial Analysis Benchmarks
#43Re: GLM-5.3 Artificial Analysis Benchmarks
#44I understand that running these benchmarks can get expensive, but it would be really nice to see AA include more benchmarks of models at reasoning settings other than the maximum, at least for the biggest releases. They have that nice graph of cost vs. composite benchmark score with the Pareto frontier line, but who knows if those are actually the optimal choices? There are already a few non-max-reasoning models on t…
Re: GLM-5.3 Artificial Analysis Benchmarks
#45Earlier quoted context omitted.
Muse Spark has a nice balance. not to mentions the Contribs version is old deepseek flash prices.
Tested muse spark 1.2 because it was rated so high on design arena, and I've missed a model that can do nice UI in the hands of an operator with no UI skills. It produced worse UI mockups than GPT and GPT models are already the bottom of the barrel here. The only model that performed well was Kimi K3 - insanely good, but expensive. It's hard to trust benchmarks these days.
If you want to generate a UI based on specific user input of some kind, then they are.
I'd suggest using one model for UI and another model for tacking onto that UI. LLMs are great at pattern matching, and benchmarks don't really capture one-shotting desirable UI.
That said, benchmaxxing is a thing and your experience with models is a thing. Benchmarks are fuzzy and should be taken with a grain of salt.
Re: GLM-5.3 Artificial Analysis Benchmarks
#46Re: GLM-5.3 Artificial Analysis Benchmarks
#47Earlier quoted context omitted.
I use the $200 plan w/ Anthropic and run out of tokens half way through the week and supposedly they are progressively reducing the limits on all their subs even further. At some point I will switch, $200 buys a lot of tokens on OpenRouter.
Is the conventional wisdom that the subscription price/token is better than the API price/token not valid any more? Or is access to model diversity worth the increased per token costs?
Re: GLM-5.3 Artificial Analysis Benchmarks
#48Re: GLM-5.3 Artificial Analysis Benchmarks
#49Is it worth using these models if I have a claude code subscription already? The appeal of lower cost is nice but I haven't gotten over the switching cost yet.
Re: GLM-5.3 Artificial Analysis Benchmarks
#50And reminder: it's less than a quarter the size of Kimi K3!