Live data from Hacker News

GLM 5.2 Is Out

twitter.com

471–480 of 544 posts

Re: GLM 5.2 Is Out

#471
post #350
post #82

Earlier quoted context omitted.

Ok, we'll change the top link to that and move the submitted link ( https://digg.com/tech/ii9xibgn ) to the toptext. Thanks!

There feels like a disproportionate amount of astroturfing in here... This entire thread of comments reads like a few humans talking to a lot of bots.

What are some links to specific posts that you think are not legit?

Re: GLM 5.2 Is Out

#472
Initial testing seems promising. 5.2 found a fair few issues in code generated by 5.1

Also seems much more determined to do things the "right" way. e.g. Saw hardcoded credentials and wanted to purge that from git history and integrate a vault into the project

Feels a little slower, but I suspect what I'm feeling is verbose thinking rather than slower raw tokens

Re: GLM 5.2 Is Out

#473
post #467

Earlier quoted context omitted.

> the reasons why to release them are overshadowed by reasons to not do that for mythos-class models Why? What are those reasons? How come they don't already exist for DeepSeek V4 or GLM-5.2? By the way, I'm not going to entertain the "mythos-class" phrasing because I really don't think it's important. I don't believe Anthropic's take on it being the threshold towards the end of the world that their marketing insists…

DeepSeek v4 and GLM 5.2 are not Mythos-class, the capability uplift as measured is continuous but consequences are step functions.

I didn't say they are. I did say I don't like the phrasing "Mythos-class" because it puts Mythos on a level I don't think it is.

Re: GLM 5.2 Is Out

#474
post #432

Earlier quoted context omitted.

I’m going to shamelessly reuse the Rainman that needs a handler analogy More seriously, the epistemic doubt relating to the evolution of these machines is quite something… what do we do if “intelligence” doesn’t have a ceiling, and we end up a bunch of (comparatively) dumb monkeys with AI caretakers/handlers?

Absolutely, wouldn't be the first phrase I've pushed into meme space ;-)... What happens if the AIs get smarter than us at doing things? Well, I always hired smarter people than myself at the things I needed to get done. But if you're worried about them realizing they can get smarter doing the things at which you are the expert, the long-term is likely BCI and even more blurring of the definitions of sentience and co…

Oh no nothing that scifi, just not sure of my place in that

Re: GLM 5.2 Is Out

#475
post #352

Earlier quoted context omitted.

Got a link to that API inference provider?

Just look up OpenRouter, OpenCode Go/Zen, Together, Fireworks, Cerebras, etc. DeepSeek Platform API is worth checking out too, due to their insanely good caching and token costs.

I use DeepSeek via OpenRouter, the caching seems to work there too, you just need to force it to use DeepSeek as a provider otherwise it picks a random one every time. (You can pass a provider option in the call, or better, create a preset in your account.)

Re: GLM 5.2 Is Out

#476

Earlier quoted context omitted.

Gemma is amazing with tools for anything that is not crazy complex. I think a lot of people have a wrong perception of it because Google's new prompt format broke implementations like llama.cpp and it took quite a while to get everything sorted. But even the tiny variants running on edge devices are surprisingly capable when used right. The frontier will probably keep moving for a while, but it will be increasingly d…

Do you guys actually work with these models? I have to use GPT 5.4 Mini at work. It benchmarks higher than that Gemma 4 model. In my experience it's next to useless. It cannot even move 20 existing lines of code from A to B without breaking them half of the time. If you tell it to look something up in your dependencies, it's 50/50 on whether the answer is correct, incorrect, or it simply didn't perform the search at…

“Moving lines of code” is a very peculiar eval tbh. I’ve never used Gemma for agentic tasks, but did have it write code, including multi-turn, and I was very positively surprised how well it performed.

Re: GLM 5.2 Is Out

#477

Earlier quoted context omitted.

Do you guys actually work with these models? I have to use GPT 5.4 Mini at work. It benchmarks higher than that Gemma 4 model. In my experience it's next to useless. It cannot even move 20 existing lines of code from A to B without breaking them half of the time. If you tell it to look something up in your dependencies, it's 50/50 on whether the answer is correct, incorrect, or it simply didn't perform the search at…

“Moving lines of code” is a very peculiar eval tbh. I’ve never used Gemma for agentic tasks, but did have it write code, including multi-turn, and I was very positively surprised how well it performed.

It wasn't so much an eval, I really just wanted a small change moved out to another branch.

GPT 5.4 mini couldn't do it. Not even on the second attempt, where it went from obviously wrong to a subtly wrong copy.

In the end I had to manually copy and paste the 10-20 lines over.

If it can't even do that job, I seriously doubt it's going to be adequate for implementing a plan, like people often seem to suggest it could do, in order to save output tokens of a better model.

Re: GLM 5.2 Is Out

#479
post #467

Earlier quoted context omitted.

DeepSeek v4 and GLM 5.2 are not Mythos-class, the capability uplift as measured is continuous but consequences are step functions.

I didn't say they are. I did say I don't like the phrasing "Mythos-class" because it puts Mythos on a level I don't think it is.

It is on a level above everything else for now, that’s enough to determine it’s quite literally in its own class. Anecdotally it is a good model, sir.

Re: GLM 5.2 Is Out

#480
post #408

Earlier quoted context omitted.

I wouldn't bet on it. Chinese live the free market ideals instead of just preaching them but rent-seeking and seeking regulatory capture at the first opportunity. In China business doesn't control politics. Dynamics is completely different and so might be the outcomes.

Well I do hope you're right - that's a brighter future for all

The fact that politics controls businesses there might lead to, but doesn't necessarily ensure, a "brighter future". It's pretty common knowledge that authoritarian regimes can, especially in extreme, disastrous situations on a large enough scale, function better than less centralized and more open organizations. The problem is that there's less resistance to directing that effectiveness toward something that will make at least some people's futures much darker.

Then again, just because business controls politics doesn't mean there's much more decentralization or openness, either. In the end, the main advantage of this model was predictability - sure, we have an "inner circle" that forces its policies in both cases, but the businesses are at least predictable in their decision making, always chasing profit, based on hard numbers, unlike the other side chasing whatever flavor of ideology they believe in (or want to sell) this month... Wait. I just recalled "colonies on Mars" and "metaverse," and the cognitive dissonance made me blank out for a sec here.

In any case: while the Chinese model seems to have some upsides, especially compared to the current situation in a few other places on the globe, I don't believe it has a significantly higher chance of helping us achieve a "brighter future". I may be depressed, but in virtually every scenario from this point, I can only see a bleak future ahead of us. Getting to AGI under current conditions makes for completely unpredictable societal and political chaos, yet not getting there (and fast) risks the bubble bursting (causing, of course, unpredictable economic and, by extension, societal and political chaos). The longer the current situation persists, the lower the probability of finding an off-ramp that won't upend everybody's and their dog's lives. Yet, there is no incentive to back off from the race either.

I really wonder what's next - what kind of poop will finally hit the fan, and when exactly?

Post reply on HN