For web dev is just a must to have, and offloading that part to a secondary model doesn't work really well in my experience.
GLM-5.3 Artificial Analysis Benchmarks
51–60 of 64 posts
Re: GLM-5.3 Artificial Analysis Benchmarks
#52I understand that running these benchmarks can get expensive, but it would be really nice to see AA include more benchmarks of models at reasoning settings other than the maximum, at least for the biggest releases. They have that nice graph of cost vs. composite benchmark score with the Pareto frontier line, but who knows if those are actually the optimal choices? There are already a few non-max-reasoning models on t…
You can turn on various levels of some of many of the models in the UI
Re: GLM-5.3 Artificial Analysis Benchmarks
#53...do I take out a double mortgage to buy a 4 Spark cluster?
Re: GLM-5.3 Artificial Analysis Benchmarks
#54Is it worth using these models if I have a claude code subscription already? The appeal of lower cost is nice but I haven't gotten over the switching cost yet.
no, at subscription prices claude is a better value than GLM. They're only a better value if you're paying API rates
Re: GLM-5.3 Artificial Analysis Benchmarks
#55Re: GLM-5.3 Artificial Analysis Benchmarks
#56Is it worth using these models if I have a claude code subscription already? The appeal of lower cost is nice but I haven't gotten over the switching cost yet.
FYI, you can use your Claude subscription pricing with OpenCode via Meridian[0], which also makes it easier to try out other models when they come out. You can also use your other subscriptions in OpenCode with CLIProxyAPI[1]. The switching cost was relatively high, mostly from claude code plugins but completely worth it. I'm now mostly using GLM-5.3 and Codex models via OpenCode and barely using Claude which seemed…
Re: GLM-5.3 Artificial Analysis Benchmarks
#57Re: GLM-5.3 Artificial Analysis Benchmarks
#58I've tested GLM 5.3 on the release day and Artificial Analysis is spot on. It's a really good model. But my main takeaway was something else. I've used closed weight models for long enough that I've forgotten how good it feels to see reasoning tokens. With GPT/Claude, you kind of hope that intent was captured well, that agent had all the information, all the tools it needed, because you won't see "hmmm it seems like…