Live data from Hacker News

GLM 5.2 Is Out

twitter.com

511–520 of 544 posts

Re: GLM 5.2 Is Out

#511
post #242

Earlier quoted context omitted.

> Opus in January So pre-nerf Opus?

Was going to say, I don't think Opus has really got much better in the last 6mo. It just goes in cycles of being better and then being worse again, presumably based on how much Anthropic are having to optimise inference

[deleted]

Re: GLM 5.2 Is Out

#512

Earlier quoted context omitted.

I'm actually interested in doing that. What would be the most favorable model/company to move to for scientific programming and engineering questions?

I'd suggest using OpenCode (via Go sub or just API credits). It will give you access to more than just one companies models and you can experiment and find one that works best for you. I really like GLM and ended up subbing to both OpenCode Go & z.ai. Mistral, Kimi and Mimi are all also options as well. I have been eyeballing the Kimi Pro sub for a while now and contemplating cancelling my ChatGPT sub for it.

OpenCode Go is pretty good in my experience too.

I ended up using DeepSeek V4 Flash as main workload model, while keeping DeepSeek V4 Pro and Qwen 3.7 Plus as advisors on system architecture and other advanced matters to guide DS Flash.

I run a simple benchmark on OpenCode Go models while ago, if anyone want to read more: https://arizenai.com/seven-models-judged-each-other/

Re: GLM 5.2 Is Out

#513

Earlier quoted context omitted.

Speaking practically your hypothetical is a scenario that requires somebody that is proactively interested in, and theoretically capable of, making a e.g. dangerous virus, yet are unwilling/unable to do so without a chatbot. How many people might this possibly apply to? I think the number is literally zero. I also don't entirely understand your comment, because your latter parts do not follow from your lead. You're 1…

>Speaking practically your hypothetical is a scenario that requires somebody that is proactively interested in, and theoretically capable of, making a e.g. dangerous virus, yet are unwilling/unable to do so without a chatbot. How many people might this possibly apply to? I think the number is literally zero. I don't disagree with the rest of your post, but this doesn't seem correct. I think I'd phrase it that there p…

The important part is being theoretically capable of. Fortunately there are massive barriers to doing things like synthesizing deadly viruses, and it's not just a matter of knowledge but of skill. For instance there was a Japanese death cult [1] that at its peak included not only many graduates of top universities in Japan but tens of millions of dollars in funding. But their escapades read a lot like a satire of incompetence.

That's not to say they were harmless - they managed to kill numerous people, but they'd have killed vastly more if they just drove some trucks into crowds as is becoming a typical weapon of terrorists. And I think the main reason is because knowing how something is done, and actually doing that thing, are radically different.

For a goofy analog, think about assembling sofas or even certain desks/chairs from a kit. That can actually be fairly tricky, to the point that there's an industry built around doing it for you. But there it's literally following like a few dozen steps with a carefully manufactured set of goodies and all tools right in front of you. Imagine doing something many orders of magnitude more complex where you're improving everything, have guidance that may be simply wrong, requires not only extreme skill but also a wide variety of difficult to acquire equipment, and if you make any mistake - you stand a decent chance of killing yourself.

[1] - https://en.wikipedia.org/wiki/Aum_Shinrikyo

Re: GLM 5.2 Is Out

#516

Earlier quoted context omitted.

Seriously? What are you, a CCP spokesperson? Murder, torture, destruction of temples and trying to abolish their religion and identity? Get out.

I'm just calling it like it is. When you define "evil" to mean "political things I disagree with" then you can arbitrarily label anything as evil.

[deleted]

Re: GLM 5.2 Is Out

#517

I'm interested in seeing how this changes folks' workflows. For me, at work I use opus to plan, brainstorm, grill, ask questions about my codebase, etc. It is pretty good about understanding the codebase holistically and providing architecturally clean solutions that actually work. Then I use sonnet as a plan executor and it does well. Follows instructions and runs tests and just overall does great. At home I make so…

I've found the prompting needs are drastically different from the latest frontier models to the latest open weight models. I can be much more vague and talk about an end goal with the frontier models vs needing to be more prescriptive + have a workflow on the open weight models. This gap continues to close, but the level of abstraction I'm working on with the latest models continues to move much higher.

Re: GLM 5.2 Is Out

#518

Earlier quoted context omitted.

Do you guys actually work with these models? I have to use GPT 5.4 Mini at work. It benchmarks higher than that Gemma 4 model. In my experience it's next to useless. It cannot even move 20 existing lines of code from A to B without breaking them half of the time. If you tell it to look something up in your dependencies, it's 50/50 on whether the answer is correct, incorrect, or it simply didn't perform the search at…

Like I said in my original comment, it’s fine for non-coding tasks, meaning I primarily use it to answer questions

The MoE variant was perfect for speedily generating hundreds of vocabulary mnemonic flash cards for my daughter to study for the SAT. "Ant bait abates our ant problem" and "A droid adroitly fixes things around the house," for example.

We also used z-image to generate accompanying illustrations.

Re: GLM 5.2 Is Out

#519
post #499

Earlier quoted context omitted.

It doesn't seem to be on a level above everything else, no. It seems to be a step increase in some areas and maybe even a decrease in others. Anectodally, DeepSeek V4 is a very good model as well, sir. I'm not calling anything V4-class because of that.

I’ve been piloting frontier LLMs for as long as anyone outside of the labs and I just disagree. It is a tier above for some tasks (especially in my usage) and not a downgrade on anything I tried it on. This is enough for me to rank it higher; ymmv.

Fair enough!

I've only briefly tried it and it did seem quite capable for what I was doing, but not that much better than the Chinese models I've been mostly using.

In any case, this [0] seems to paint a more reasonable picture than "it's much better than anything else at everything".

[0] https://www.aisi.gov.uk/blog/our-evaluation-of-claude-mythos...

Re: GLM 5.2 Is Out

#520
post #497

Earlier quoted context omitted.

Yes, unironically claiming that and not wild at all if you're a practitioner. It doesn't become actual reasoning just because you chose to call it so. If they did reason, LLMs would not fail at ridiculously easy problems like strawberry or car wash ones. LLMs are great at search . They only emulate reasoning. They can't actually reason but they approximate it. Combine it with copious amount of computes and some searc…

If humans did actual reasoning, then why is this particular discussion failing so hard?

Having the ability to reason != never failing at logic.
Post reply on HN