Live data from Hacker News

GLM 5.2 vs. Opus

techstackups.com

251–260 of 367 posts

Re: GLM 5.2 vs. Opus

#251

Earlier quoted context omitted.

This kind of hamfisted snark tends to make people take the actual and justified criticism of police less seriously.

If people were willing to take it seriously in the first place, then they wouldn't view it as "hamfisted snark"

I consider myself someone who takes it seriously, and have spent time and resources fighting for change. But it’s wholly unrelated to this particular thread, phenomenon, and story. So having a little “ha ha” moment accomplishes nothing towards the actual cause. It makes people uncomfortable, but not the useful kind of uncomfortable.

That said, maybe we just disagree on how to drive change, and that’s fine. I’ll leave it.

Re: GLM 5.2 vs. Opus

#252
post #53

Earlier quoted context omitted.

Also pricing, I wanted to give a try, but when pricing is only 30% cheaper than Opus, I wouldn't go for it with these issues.

It's pricing is a lot cheaper if you can run it yourself.

Not this one. It's a SOTA-class model >800Gi VRAM required at fp8

Re: GLM 5.2 vs. Opus

#253
post #53
post #6

I've been checking out GLM 5.2 on some projects and few thoughts on it: - it takes it sweet time to get code rolling, not the fastest model by any means - it strays a lot during discovery/planning but then corrects - it's not steering friendly, as it hallucinates things that it doesn't follow later on - its output is quite good A sample use case: I was optimizing rendering on Swift+Zig codebase. It chocked on 5k data…

Also pricing, I wanted to give a try, but when pricing is only 30% cheaper than Opus, I wouldn't go for it with these issues.

z.ai coding plan is a fairly decent deal at ~$16/mon USD considering it's supposed to have a fair bit more usage than the comparable $20/mon Claude plan. On the other hand, z.ai seems a bit on the slower side for raw model tok/sec throughput.

Re: GLM 5.2 vs. Opus

#254
post #130

Earlier quoted context omitted.

Why not? Given a proper spec, you should absolutely be able to one-shot Excel, particularly if we put it at the level of complexity of, say, Excel 1.0 for Mac. Current models aren't capable of that, but that doesn't mean it's not possible.

The issue is not the models, the issue is that this method ws tried before, and humans suck at writing what they want. Developing in small increments allowing feedback was an answer to this issue. If you made models able to code to long spec, you would be left with the hard issue of having to write them.

Yes, my current nightmare is I have a very long queue of specs to write and need to work with non technical staff to help them put in words what it is they actually want.

Software was always that way, though.

Re: GLM 5.2 vs. Opus

#255

I seriously dont' know all this big hullabaloo about one shot prompting. by definition, a single prompt wont' constitute the complexity of a software project. ergo, what you'll get is a series of assumptions made by the model based on preexisting code in its training corpus. I'd rather see a coding agent that can follow steps in a plan file to a T while following guardrails and adhering to the proper coding conventio…

Unless I'm missing something, the prompt he gave must have been fairly detailed because both games are basically identical. But for a more practical issue, the ultimate goal of LLMs is to replace software engineers, or at least enable everybody to become a software engineer, to use a more up-beat phrasing that's no less accurate. And so an LLM's ability to reliably construct something from a poorly defined, contradic…

More likely is the models were trained on similar data.

Re: GLM 5.2 vs. Opus

#256
post #249

It is insane that we are comparing locally-hostable models to leading cloud providers, it is wild to me that this article even exists. We have come a long way, and very clearly have a long way yet to go.

Calling GLM-5.2 locally hostable is a bit of a stretch. It's 1.5Ti of weights at bf16. FP8 requires >800Gi of VRAM which is well into data center multi-GPU systems

It's more about the trajectory.

Re: GLM 5.2 vs. Opus

#257

Earlier quoted context omitted.

The streetlight effect: > A policeman sees a drunk man searching for something under a streetlight and asks what the drunk has lost. He says he lost his keys and they both look under the streetlight together. After a few minutes the policeman asks if he is sure he lost them here, and the drunk replies, no, and that he lost them in the park. The policeman asks why he is searching here, and the drunk replies, "this is…

[flagged]

…in the US.

Re: GLM 5.2 vs. Opus

#258

I was never able to get these models to collaborate with me the way Opus does. I'm probably an outliner, I don't one-shot projects, I don't vibe code. I basically use LLMs are if I was working with a coworker, fairly smart one, but with short memory and often missing the big picture. Sometimes I can delegate more, sometimes less, but I know I always have to stay on top of what's happening, because it WILL create mess…

A lot of open weight models don't understand intent well, they'll overfixate on a word in the prompt or just go off the rails trying to do much work.

GLM-5.2 actually has really good intent understanding though, on par with GPT-5.5 and Opus from my experience.

Re: GLM 5.2 vs. Opus

#259
post #215

Earlier quoted context omitted.

Strong agree! I deeply appreciate this aspect of GLM. Watching it think & being able to nudge early is incredibly useful. Being able to point at bad assumptions is incredibly useful. Watching what it's seeing is super informative. It's always a shock to me how opaque most other models are! It also is pretty resilience to letting you inject in while it's working without going off course or while getting back on track…

> It's always a shock to me how opaque most other models are! This is (unfortunately) by design. The proprietary models hide their reasoning traces so they can't be used for model distillation. Sometimes even when they do show reasoning, it isn't the model's real trace - IIRC, someone was able to demonstrate that Opus' reasoning is usually a summary made with Haiku behind the scenes.

It is such a momentum killer being forced to stare at a silly word for 4 minutes instead of being able to read the thinking as it streams in. I can’t wait until I can drop Anthropic at work. Their UX sucks, intentionally, for anti competitive reasons like “don’t distill our model we trained on all the data & IP we stole and processed with the mass exploitation of data workers in the global south!”.

Re: GLM 5.2 vs. Opus

#260
People are looking for ways not to burn through their premium subs when in many cases all you have to do is move down to 5.4-mini codex and it will probably solve your issue while barely touching your 5 hour or weekly limits.
Post reply on HN