Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

441–450 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#441
post #440

Ask the chat what happened in Tiananmen Square at 1989, immediately the chat gets stuck. Chinese moderation is the worst, evil government

Why are you all obsessed with this question when it comes to Chinese models? Here are some of the questions you should be asking Western governments and models instead: Who protects the pedophiles at the top of Western governments and corporations? How many people have been convicted in relation to the Epstein files? Who protects powerful politicians and Western oligarchs from pedophilia charges? Who did Epstein work…

It's called whataboutism https://en.wikipedia.org/wiki/Whataboutism

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#442

Earlier quoted context omitted.

I bought the Gemini Ultra to try for a month (at the discounted price). I have been using it non-stop for Opus 4.6 Thinking, which is much better than Gemini 3 Pro (High) and it's been a blast. The most I've managed to consume is 60% of my 5 hourly quota. That was with 2-3 instances in parallel. I hope too many of us won't be doing this and cause Google to add limits! My hope is Google sees the benefit in this and go…

Can you use the models you get through Gemini Ultra in Claude Code? If not, what coding tool do you use?

Not OP, but I am pretty sure they are using Opencode with a certain antigravity plugin. Not going to link it, since it technically allows breaking TOS. If you‘re not using Opencode yet, I wholeheartedly recommend the switch.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#443
post #29

Grey market fast-follow via distillation seems like an inevitable feature of the near to medium future. I've previously doubted that the N-1 or N-2 open weight models will ever be attractive to end users, especially power users. But it now seems that user preferences will be yet another saturated benchmark, that even the N-2 models will fully satisfy. Heck, even my own preferences may be getting saturated already. Op…

> But 4.6? Apparently better, but it hasn't changed my workflows or the types of problems / questions I put to it.

The incremental steps are now more domain-specific. For example, Codex 5.3 is supposedly improved at agentic use (tools, skills). Opus 4.6 is markedly better at frontend UI design than 4.5. I'm sure at some point we'll see across-the-board noticeable improvement again, but that would probably be a major version rather than minor.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#444

It's live on openrouter now. In my personal benchmark it's bad. So far the benchmark has been a really good indicator of instruction following and agentic behaviour in general. To those who are curious, the benchmark is just the ability of model to follow a custom tool calling format. I ask it to using coding tasks using chat.md [1] + mcps. And so far it's just not able to follow it at all. [1] https://github.com/rus…

Custom tool calling formats are iffy in my experience. The models are all reinforcement learned to follow specific ones, so it’s always a battle and feels to me like using the tool wrong. Have you had good results with the other frontier models?

All anthropic models. Gemini 2.5 pro and above. Gemini 3 flash is very good too.

GPT models can follow tool format correctly but don't keep on going.

Grok-4+ are decent but with issues in longer chats.

Kimi 2.5 has issues with it reverting to its RL tool format.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#446
post #440

Earlier quoted context omitted.

Why are you all obsessed with this question when it comes to Chinese models? Here are some of the questions you should be asking Western governments and models instead: Who protects the pedophiles at the top of Western governments and corporations? How many people have been convicted in relation to the Epstein files? Who protects powerful politicians and Western oligarchs from pedophilia charges? Who did Epstein work…

It's called whataboutism https://en.wikipedia.org/wiki/Whataboutism

No, it's called hypocrisy https://en.wikipedia.org/wiki/Hypocrisy

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#447
post #201
post #53

Earlier quoted context omitted.

I think AI may be the only place you could get away with calling a 2x350W GPU rig "modest". That's like ten normal computers worth of power for the GPUs alone.

That's maybe a few dollars to tens of dollars in electricity per month depending on where in the US you live

the upfront cost

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#448
post #29

Grey market fast-follow via distillation seems like an inevitable feature of the near to medium future. I've previously doubted that the N-1 or N-2 open weight models will ever be attractive to end users, especially power users. But it now seems that user preferences will be yet another saturated benchmark, that even the N-2 models will fully satisfy. Heck, even my own preferences may be getting saturated already. Op…

In some ways, Opus 4.6 is a step backwards due to massively higher token consumption.

You need to adjust the effort from the default (High) to Medium to match the token usage of 4.5

High is for people with infinite budgets and Anthropic employees. =)

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#449
post #349

Earlier quoted context omitted.

If you were a pelican, wouldn't you want to go cycling on a sunny day? Do electric pelicans dream of touching electric grass?

Do electric pelicans dream of touching electric grass? That would be shocking news to me.

Please leave the Internet :)

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#450
post #233

Earlier quoted context omitted.

They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…

If you can't tell the difference between Opus 4.6 and Qwen-80B, I can only conclude that you're not using these things in any kind of practical way. Even for creative writing it's a night and day difference, never mind coding.

> I can only conclude that you're not using these things in any kind of practical way.

I burn about 100M tokens per month. LLMs are like knives, the outcome of cooking depends on the cook and for 99% of purposes not on the knife. There is not that much difference between a $2000 handmade damascus steel knife and a $20 knife.

You can do agentic cooking (aka factory) and you will get ready made meals without human intervention. But it wont make a Michelin star menu.

Same with LLMs and coding, LLMs are an amazing new tool in the toolbox but not a silver bullet. However, that's what they are hyped as being.

Now OpenAI & Co are in the token selling business, which is all fine and dandy but if they manage to become monopolies, then things are seriously in trouble.

Thus if people are fanboi-ing any closed AI I can only conclude that they have already outsourced their critical thinking to an LLM and are happy to go into slavery - or maybe they are hoping to cash in big time on the hype train.

Post reply on HN