Live data from Hacker News

GLM 5.2 Is Out

twitter.com

501–510 of 544 posts

Re: GLM 5.2 Is Out

#501

Earlier quoted context omitted.

“Moving lines of code” is a very peculiar eval tbh. I’ve never used Gemma for agentic tasks, but did have it write code, including multi-turn, and I was very positively surprised how well it performed.

It wasn't so much an eval, I really just wanted a small change moved out to another branch. GPT 5.4 mini couldn't do it. Not even on the second attempt, where it went from obviously wrong to a subtly wrong copy. In the end I had to manually copy and paste the 10-20 lines over. If it can't even do that job, I seriously doubt it's going to be adequate for implementing a plan, like people often seem to suggest it could…

Like I said, I never really used it for agentic work. I had previously evaluated locally runnable models with opencode (such as qwen3-coder), but found that it wasn't really feasible.

Since then I've adopted a different philosophy, and I actually prefer it this way.

I still very much enjoy doing most coding myself, but when I tried using tools like Claude Code, it felt very difficult to return to the codebase after letting Claude make some changes. Maybe that's just because of poor AI-use discipline, I don't know. But with smaller models, that's not even an issue. I can't just let it do all the coding and thinking for me, however if I can describe a function I want to great detail in plain english, then Gemma can write it for me, and it will most likely work. It's perfect for boilerplate.

I also recently worked with a web framework I'd never worked before, though I'm deeply familiar with other ones. So I asked it "I know how to do this in Y framework, what's the best-practice approach to doing it in Z framework?" and it was incredibly helpful, even pushing back on some of my 'bad' attempts at solving a problem.

I think GPT5.4 mini might fall into a similar category, in that it probably performs best when not overwhelmed with too many tools/ skills/ mcps, instead being given clearly defined tasks by an orchestrator model. I call those my token burners, as they're super cheap to run and have high tokens/second.

Re: GLM 5.2 Is Out

#502
post #479

Earlier quoted context omitted.

It is on a level above everything else for now, that’s enough to determine it’s quite literally in its own class. Anecdotally it is a good model, sir.

It doesn't seem to be on a level above everything else, no. It seems to be a step increase in some areas and maybe even a decrease in others. Anectodally, DeepSeek V4 is a very good model as well, sir. I'm not calling anything V4-class because of that.

Have you used it? It’s clearly a class above, I had it solve so many things in 3 days, it was ridiculous

Re: GLM 5.2 Is Out

#503
post #238

Earlier quoted context omitted.

I think that this is what OpenAI/Anthropic want but they wont say it publicly. The will be OK with the US banning regulating and banning open source models as it let's Anthropic and OpenAI charge huge premiums to American business clients for their models. Also the marketing of them getting to say "our models are so dangerous" only a few companies or select users are allowed to use (benchmark) them would help keep th…

> I think that this is what OpenAI/Anthropic want but they wont say it publicly. Won't say it publicly? Anthropic is openly and explicitly saying it publicly. Here: https://darioamodei.com/post/policy-on-the-ai-exponential > AI companies that develop advanced AI models must have strong security standards that protect their model weights If the model is open-weight then there's nothing to protect, so the only way to f…

> Won't say it publicly? Anthropic is openly and explicitly saying it publicly. Here: https://darioamodei.com/post/policy-on-the-ai-exponential

Off-topic, but tech-bros fixation on LotR (benevolent[0] or not[1]) makes me sick to my stomach.

[0] https://lucumr.pocoo.org/2026/1/27/earendil/

[1] https://en.wikipedia.org/wiki/Palantir , https://en.wikipedia.org/wiki/Mithril_Capital , https://en.wikipedia.org/wiki/Anduril_Industries

Re: GLM 5.2 Is Out

#504
post #472

Initial testing seems promising. 5.2 found a fair few issues in code generated by 5.1 Also seems much more determined to do things the "right" way. e.g. Saw hardcoded credentials and wanted to purge that from git history and integrate a vault into the project Feels a little slower, but I suspect what I'm feeling is verbose thinking rather than slower raw tokens

I wonder how many issues 5.1 could have caught if you ran it as multiple adversarial reviews against the original output it gave you.

Well I did run deepseek against the original. They all seem seem to spot different issues

Re: GLM 5.2 Is Out

#505
post #421

Earlier quoted context omitted.

Restricting things like creation of a highly infectious virus is very different from restricting books or even guns. There is no 'monopoly' over such a technology, as a use of the technology will inevitably harm the creators themselves. Restrictions on high end biology, chemistry would leave overwhelming number of use cases of LLMs unaffected - no need to ban open weight LLMs. Such restrictions can be even more effec…

Speaking practically your hypothetical is a scenario that requires somebody that is proactively interested in, and theoretically capable of, making a e.g. dangerous virus, yet are unwilling/unable to do so without a chatbot. How many people might this possibly apply to? I think the number is literally zero. I also don't entirely understand your comment, because your latter parts do not follow from your lead. You're 1…

>Speaking practically your hypothetical is a scenario that requires somebody that is proactively interested in, and theoretically capable of, making a e.g. dangerous virus, yet are unwilling/unable to do so without a chatbot. How many people might this possibly apply to? I think the number is literally zero.

I don't disagree with the rest of your post, but this doesn't seem correct.

I think I'd phrase it that there probably already exist, or will exist, people with the inclination to cause global mass death, but don't have the knowledge or ability to manufacture a virus to achieve this.

Re: GLM 5.2 Is Out

#506

Earlier quoted context omitted.

Are you unironically claiming that LLM's can't reason? That's an absolutely wild claim in an era where they're solving Erdos problems and writing better code than many senior devs. What's the basis for it? Agency is harder to define, but most any definition I can come up with LLM's meet. Again, I'm curious how you define it in a way that excludes frontier models but doesn't also exclude many humans.

Yes, unironically claiming that and not wild at all if you're a practitioner. It doesn't become actual reasoning just because you chose to call it so. If they did reason, LLMs would not fail at ridiculously easy problems like strawberry or car wash ones. LLMs are great at search . They only emulate reasoning. They can't actually reason but they approximate it. Combine it with copious amount of computes and some searc…

> They only emulate reasoning.

So do we.

Re: GLM 5.2 Is Out

#507
post #421
post #410

Earlier quoted context omitted.

The printing press gave us the renaissance, even though the church argued it was too dangerous to give non-clergy access to books. Even things like universal access to guns was a net positive. It led to the end of feudalism and rise of democracy. The sad truth is that whenever any one group of people gets a monopoly over an important technology, they use it to exploit/enslave/murder everyone they can. Look at the int…

Restricting things like creation of a highly infectious virus is very different from restricting books or even guns. There is no 'monopoly' over such a technology, as a use of the technology will inevitably harm the creators themselves. Restrictions on high end biology, chemistry would leave overwhelming number of use cases of LLMs unaffected - no need to ban open weight LLMs. Such restrictions can be even more effec…

I'm amazed we didn't have the same moral panic when the web became popular. billions of people suddenly had access to knowledge about how to create dangerous viruses! sites like Wikipedia don't even check that you're a US citizen before letting you access pages on recombinant DNA and genetic engineering! the articles on sarin and VX nerve gas include syntheses!

Re: GLM 5.2 Is Out

#508
post #421

Earlier quoted context omitted.

Restricting things like creation of a highly infectious virus is very different from restricting books or even guns. There is no 'monopoly' over such a technology, as a use of the technology will inevitably harm the creators themselves. Restrictions on high end biology, chemistry would leave overwhelming number of use cases of LLMs unaffected - no need to ban open weight LLMs. Such restrictions can be even more effec…

I'm amazed we didn't have the same moral panic when the web became popular. billions of people suddenly had access to knowledge about how to create dangerous viruses! sites like Wikipedia don't even check that you're a US citizen before letting you access pages on recombinant DNA and genetic engineering! the articles on sarin and VX nerve gas include syntheses!

Wikipedia is a presentation of partial selection of biology textbooks and research papers, not using them as a collective brain to generate new artifacts.

There is a big difference between having a large bookshelf of programming language/networking/OS manuals and the ability to generate a functional software product which previously required a hundred or more developers. Even a hundred developers may not be able to find a subtle exploit in code which requires a tedious scan of millions of lines. Computer security hacks can be much less of a problem in comparison to exploits in biology.

Also, even Wikipedia (and public resources in general) have restrictions - there is information dangerous enough to be not published. In the 1930's itself, Szilard (who discovered the chain reaction) and Bohr advocated for restrictions on openly publishing research on uranium fission.

Re: GLM 5.2 Is Out

#509

Earlier quoted context omitted.

Yes, unironically claiming that and not wild at all if you're a practitioner. It doesn't become actual reasoning just because you chose to call it so. If they did reason, LLMs would not fail at ridiculously easy problems like strawberry or car wash ones. LLMs are great at search . They only emulate reasoning. They can't actually reason but they approximate it. Combine it with copious amount of computes and some searc…

> They only emulate reasoning. If they emulate reasoning well enough that it gets the same or better results what is the difference? Semantics? I can't help but wonder if you dont percieve what they do as reasoning because its different from the way you reason? > strawberry or car wash ones. Humans fall for the Nigerian scam still. We all have blind spots but that doesnt imply we're all completely blind.

I run Claude Max daily, and tried letting Opus 4.8 write an ADR with known requirements.

After searching through codebase, git history, etc it spat out a surface level reasonable ADR, with the customary bloated text.

I started reading through it asking "Is this sentence needed?: ''", whereby it acknowledges that no, it adds nothing and changes nothing not already served by other statements. I ask it to go through each sentence one by one asking the same question. It claims to do so, and give me two suggestions to remove in the entire document.

I then spend a few more minutes giving 10 additional sentences manually that it happily acknowledges are redundant.

I ask why those weren't removed in my previous prompt, and frankly I can't remember specifically what rationalization it gave, I assume because it's not memorable because there can be none, because it very obviously is not reasoning.

Re: GLM 5.2 Is Out

#510
post #421

Earlier quoted context omitted.

Restricting things like creation of a highly infectious virus is very different from restricting books or even guns. There is no 'monopoly' over such a technology, as a use of the technology will inevitably harm the creators themselves. Restrictions on high end biology, chemistry would leave overwhelming number of use cases of LLMs unaffected - no need to ban open weight LLMs. Such restrictions can be even more effec…

Speaking practically your hypothetical is a scenario that requires somebody that is proactively interested in, and theoretically capable of, making a e.g. dangerous virus, yet are unwilling/unable to do so without a chatbot. How many people might this possibly apply to? I think the number is literally zero. I also don't entirely understand your comment, because your latter parts do not follow from your lead. You're 1…

If it just a mundane chatbot, the discussion is moot. But, we already have AI making breakthroughs in research and approaching the abilities do science just like a scientist does. (The last two paragraphs of your comment also assume such a high capability scenario).

Imagine giving the access, to whoever wants it, to a scientist who may not have many fresh insights, but has the advantage of a huge memory containing all the scientific literature in their mind, the standard patterns of deductions, and the ability to work at a very fast pace 24/7. They could identify vulnerabilities in biological mechanisms, just like AI identifies security flaws in code today.

---

Regarding hurting themselves, I was not referring to someone who is too dumb to follow lab safety precautions, but someone who has a nihilistic mindset. State actors and militia use weapons to take over and enjoy the power they acquire - they dont want to get killed by a deadly virus(unless they engineer and selectively apply the vaccine before they release the weapon - but this is very hard to keep secret). Someone who is nihilistic wont have such reservations on using the weapon even if it destroys them eventually.

Regarding restrictions on API LLMs leading to use of local LLMs, it is the local LLMs which will be used anyway (once they have the capability). That we live in a mass surveillance envirnoment is common knowledge. The bottleneck, where restrictions can be applied, is not inference but training which requires hundreds of millions of dollars. Chinese scientists have themselves spoken about AI safety concerns and it is indeed a threat to China just like anyone else.

Also, restricting high end weapons ability does not interfere with 99.9% of LLM usage (open-weights or proprietary) - so it need not interfere with business strategy.

Post reply on HN