Live data from Hacker News

GLM-5.3 Artificial Analysis Benchmarks

artificialanalysis.ai

11–20 of 64 posts

Re: GLM-5.3 Artificial Analysis Benchmarks

#13
post #7

Is it worth using these models if I have a claude code subscription already? The appeal of lower cost is nice but I haven't gotten over the switching cost yet.

Use a unified proxy that lets you switch between models seamlessly. We are far from an equilibrium in this market and you will continue to have FOMO no matter who you pick if you go all in on one company

Re: GLM-5.3 Artificial Analysis Benchmarks

#14
post #8

Very impressive score for the size, though token use is higher than k3 and far higher than proprietary models, and its price to performance isn't all that far ahead of k3 as a result

>token use is higher than k3 and far higher than proprietary models GLM sets effort to max by default historically.

Aa also benchmarked k3 at max

Re: GLM-5.3 Artificial Analysis Benchmarks

#15
post #7

Is it worth using these models if I have a claude code subscription already? The appeal of lower cost is nice but I haven't gotten over the switching cost yet.

no, at subscription prices claude is a better value than GLM.

They're only a better value if you're paying API rates

Re: GLM-5.3 Artificial Analysis Benchmarks

#16

I like to compare models with a similar score on cost per task and output tokens per task since those measure two things I'm interested in: cost efficiency and token efficiency. Here's how GLM-5.3 compares to other models in a similar score and against GLM-5.2 to save a few clicks for others who care about these metrics: Model Score Cost / Task Output Tokens / Task ----------------------------------------------------…

Muse Spark has a nice balance. not to mentions the Contribs version is old deepseek flash prices.

Re: GLM-5.3 Artificial Analysis Benchmarks

#17

I like to compare models with a similar score on cost per task and output tokens per task since those measure two things I'm interested in: cost efficiency and token efficiency. Here's how GLM-5.3 compares to other models in a similar score and against GLM-5.2 to save a few clicks for others who care about these metrics: Model Score Cost / Task Output Tokens / Task ----------------------------------------------------…

Muse Spark has a nice balance. not to mentions the Contribs version is old deepseek flash prices.

I found the sweetspot here: GPT-5.6 Sol (high) 57.3 $0.52 7,545

(Edit: TLDR; It gets on with it, makes the same mistakes you would, without overthinking and overengineering, most of the time)

Re: GLM-5.3 Artificial Analysis Benchmarks

#18

I like to compare models with a similar score on cost per task and output tokens per task since those measure two things I'm interested in: cost efficiency and token efficiency. Here's how GLM-5.3 compares to other models in a similar score and against GLM-5.2 to save a few clicks for others who care about these metrics: Model Score Cost / Task Output Tokens / Task ----------------------------------------------------…

these $/task figures aren't very useful in my experience. it doesn't tell you how well it did the task.

generally I choose models by their intelligence and then personal preference from direct experience.

Re: GLM-5.3 Artificial Analysis Benchmarks

#19
post #11

Beware of the benchmarks listed. SciCode and EnterpriseOps for instance: https://shukla.io/blog/2026-08/gym.html

The Chinese models also like to cut corners on stuff like science. Their scores on stuff like biotech and scientific knowledge is far from ChatGPT unfortunately. (Claude is pretty good but it just refuses all prompts).
Post reply on HN