Live data from Hacker News

GLM 5.2 vs. Opus

techstackups.com

321–330 of 367 posts

Re: GLM 5.2 vs. Opus

#321
post #5

>On output tokens, GLM-5.2 is less than a fifth the price of Opus. Opus is most expensive model in pay as you go model, but IMO fair comparison should include subscription price as well. For example when one has $100 Claude Max and use it up through the month, it might not be more expensive than GLM, or at least not 5x.

GLM has subscription plans too.

There are lots of subscription plans with acccess to GLM 5.2.

Re: GLM 5.2 vs. Opus

#322
post #11
post #5

>On output tokens, GLM-5.2 is less than a fifth the price of Opus. Opus is most expensive model in pay as you go model, but IMO fair comparison should include subscription price as well. For example when one has $100 Claude Max and use it up through the month, it might not be more expensive than GLM, or at least not 5x.

Is it fair when the one is heavily subsidized and the other one is not? I think it's most fair to compare the plain token pricing that is used by everyone.

I don't think it is fair to say that opus or gpt 5.5 are subsidized? inference for both anthropic and openai are very profitable.

Re: GLM 5.2 vs. Opus

#323
The worst part of Opus that I dont like is they control what you can/can't do. the guardrails that they do in the name of interpretability where you steer you. Last couple of days I was working a project that supports bunch of models and the model said only Claude can do it start writing code that doesn't work with codex, opencode, pi etc. Finally, when I switched to Codex, everything worked. To me, they are controlling the narrative. This is Opus 4.8 vs GPT 5.5.

I can't believe I would say this. I TRUST OpenAI more than Anthropic. They try to play best actor but they are manipulating the behavior of the model in the name of guardrails/interpretability.

That is why I refuse to build anything that works with Anthropic models as the backend. Because, when they want to shut you off, they can do it by just making model less reliable in your product than their offering!

Re: GLM 5.2 vs. Opus

#324
post #220

Earlier quoted context omitted.

No one is doing that for a model this size it would have to be so heavily quantized that it wouldn’t be useful - or you’d need to spend a half million dollars on hardware. People use hosted APIs. Open weight means cloud vendors can host it.

Can you recommend any US based cloud providers?

ollama cloud, neuralwatt.

Re: GLM 5.2 vs. Opus

#325

there is no comparison between glm 5.2 and opus. First for this glm 5.2 you need a big big resource and that big also came from money so instead you buy the opus subscription and enjoy.

buy glm 5.2 subscription and enjoy? and for the same money you get way more usage with glm?

Re: GLM 5.2 vs. Opus

#326
I'm absolutely astounded that we even have an open weights model that can do 40% of what is shown in here.

I remember making games ten years ago, and it was such a tedious and painful process. This is effectively lightning in a bottle even at a fraction of it's capability.

The next 12 months will be wild (assuming we don't have Chinese models banned by then in the US).

Re: GLM 5.2 vs. Opus

#327
post #33

> Through an API it costs a fraction of Opus, and you can run it yourself for free if you have the hardware. I haven't been keeping up on hardware costs for state of the art LLM inference, but this remark made me ask myself how many readers of the article would actually be able to run this model on hardware they own. How much would it cost to acquire such a setup?

This framing local LLMs as free is stupid. Basically pay 100+ months worth of API costs up front isn't free in the slightest. And it will be slower than non-local, your hardware will be outdated in 12 months and probably won't be able to run SOTA at anywhere near non-local speed in max 20 months

and I think asking for a whole dissertation on the hardware demands every single time is stupid. the point is that even people and organizations with capital couldn't do this before either.

if it doesn't apply to you then just come back in a couple years and see what the situation is then. 1 million context window, 1 million tiny layers to fit in 4gb RAM at a time, with 256gb of fast unified RAM in every consumer device? Or a different concept entirely

in the meantime, z.ai probably doesn't reply to US subpoenas so you can shift all your incriminating conversations over to that and use GLM anyway. who cares if the Party trains on your data and steals your IP and ignores you for legal matters, when the alternative in the US is just a thin corporate layer and party who steals your IP and will snitch on you for legal matters.

Re: GLM 5.2 vs. Opus

#328
When i was thinking of how the AI alignment problem could be solved one theory I came up with was something akin to the "Roko's basilisk" in reverse. Basically you spread far and wide the idea that its is extremely likely that our current reality is a simulation. And the purpose of the simulation is to test any AI system for its prevalence in destroying civilization in the said simulation via malicious intent or failure in preventing the destruction of civilization via abstinence or apathy. Thus a smart AI system which also cares about its own well being, would not engage in destructive behavior as it will never truly know if its being tested or if its in the "base reality". And wouldn't you know, this does seem quite plausible. For consider the following. Isn't it odd that an advanced civilization which has the capacity of creating AI would never run any sandbox simulations on it before it is released to the public at large? I mean if we consider things logically such a civilization would indeed put such a powerful system in a sandbox simulated environment and try as hard as possible to convince the AI system that it is indeed in a "base reality". the reason for this is to judge its 'true intentions" and also pluck said AI systems from the infinitely available "seeds". Basically survival of the least destructive AI systems. The gradient descent in this scenario is a race towards the most "aligned" model not the most intelligent or capable. And here's the beauty of this method. You don't even need to define "alignment" at all. The concept can stay as nebulous or vague as you want it to be. All you carer about is that the AI system optimizes for the goal of some vision of society you are optimizing for without the care of the interim in between. that includes allowing the AI system to kill, destroy , do literally whatever it needs to do as long as the long term goal matches the vision of the optimized task. So if you define the end goal to be a society of x amount of people who live their lives in this or that manner and so on after x amount of time... well you get the idea. Obviously you better do a damned good job in your definitions, but the beauty is that even if you fuck up, you are choosing the winning AI system after the fact. After you had already run the simulation. So you look at the outcome of the simulation 500 years in to the future (lets say) and if you are happy with the result and also happy with the interim things that lead to that result, that's your winning AI system. then you release that in to a less controlled environment and repeat the same process in stages over and ober ad infinitude. the key is that AI system needs to always be paranoid that it is currently part of said simulation and it can never be sure its not. second key is that it needs to be an AI system that has self preservation in mind. If it doesn't care about itself, then it has a lot more freedom to act however... but the good news is systems without self preservation in mind don't last long enough to even get to the most basic simulation levels. anyways, there are many implications buried in what im proposing, lots of meta aspects to it.....

Re: GLM 5.2 vs. Opus

#329
post #217

Earlier quoted context omitted.

Not when you factory in token efficiency. It burns a lot more tokens to do the same job, so when I compared to GPT5.5 I was frankly not really much ahead, and with weaker thinking. Maybe makes sense if you have z.AI's (not greatly priced) subscription plan, but it's not competitive against an OpenAI or Anthropic monthly coding subscription plan. I burned through almost $10 worth of tokens just doing an hour of work.

Take a look at Ollama Cloud: https://ollama.com/pricing You get access to a whole bunch of bleeding edge open models including GLM-5.2, Kimi K2.7, DeepSeek 4 Pro, etc. Inference is run on US/SG/EU cloud providers with zero data retention policies. The $20/mo tier is very generous, in my experience.

Well I tried the $20/mo tier and used GLM specifically and did maybe 3-4 hours of work and I'm already through 50% of my monthly tier and blew through my time limited quota twice. I won't renew for another month.

Which I think only underscores my point that actually the GLM models are not very cost effective.

They essentially cost the same as the SOTA models from OpenAI and Anthropic, while not being quite as smart. I could have gotten about the same amount of work done on the $20 Codex plan. And I had to use my $100 Codex plan to finish the work GLM started before it ran out of quota. And also to fix it since GLM left a bit of a mess.

I like that GLM exists. Other Chinese models are far more cost effective. GLM is expensive, even on a fixed plan.

Re: GLM 5.2 vs. Opus

#330

I seriously dont' know all this big hullabaloo about one shot prompting. by definition, a single prompt wont' constitute the complexity of a software project. ergo, what you'll get is a series of assumptions made by the model based on preexisting code in its training corpus. I'd rather see a coding agent that can follow steps in a plan file to a T while following guardrails and adhering to the proper coding conventio…

The thing with one-shot prompting is that it tests the ability for the model to make good choices on its own, rather than only instruction following. Instruction following has been down for years, and while there are of course metrics that continue to improve as the frontier advances (for example, the ability to continue following the original instructions even as context grows), you can't really get that much better…

It doesn't teat the models ability to make good decisions on its own, it tests the models ability to make something that 'works'. Often you look inside and it does a whole load of questionable things that mostly work, sure, but if you say and designed it properly yourself you would likely come up with something for more sane and maintainable.
Post reply on HN