Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

461–470 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#461
post #208

Earlier quoted context omitted.

But they don't have the same human in the loop though.

that software is called autonomous agents, the term autonomous has nothing to do with human in the loop, it is the complete opposite.

Nothing changed since ’87. Machines still can’t be accountable and still shouldn’t make managerial decisions. Acceptance control is one of those decisions, and all the technical knowledge still matters to form a well-informed one. It may change, of course, but I have an impression that those who try otherwise seem to not fare well after the initial vibecoding honeymoon period. Of course, it varies from case to case - sometimes machines get things right, but long-term luck seems to eventually run out.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#462

GLM-4.7-Flash was the first local coding model that I felt was intelligent enough to be useful. It feels something like Claude 4.5 Haiku at a parameter size where other coding models are still getting into loops and making bewilderingly stupid tool calls. It also has very clear reasoning traces that feel like Claude, which does result in the ability to inspect its reasoning to figure out why it made certain decisions…

minimax-m.2 is close

2.5 is out now too.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#463

Earlier quoted context omitted.

They had plenty of time to update their system prompts so they don't be embarrassed. I noticed whenever such meme comes out, if you check immediately you can reproduce it yourself, but after a free hours it's already updated.

thats not how it works

And yet, I witnessed from personal experiences that such memes get fixed quickly. Whether with system prompts or some other way, I don't know, but they get fixed.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#464
post #456

Earlier quoted context omitted.

If that's what they're tuning for, that's just not what I want. So I'm glad I switched off of Anthropic. What teams of programmers need, when AI tooling is thrown into the mix, is more interaction with the codebase, not less. To build reliable systems the humans involved need to know what was built and how . I'm not looking for full automation, I'm looking for intelligence and augmentation, and I'll give my money and…

That sounds like wishful thinking. Every client I work for wants to reduce the rate at which humans need to intervene. You might not want that, but odds are your CEO does. And babysitting intermediate stages is not productive use of developer time.

And the odds are good you use the models and understand them in detail while the CEO is just buying the hype, ill informed or not.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#465

Earlier quoted context omitted.

> The U.S. Court of Appeals for the D.C. Circuit has affirmed a district court ruling that human authorship is a bedrock requirement to register a copyright, and that an artificial intelligence system cannot be deemed the author of a work for copyright purposes > The court’s decision in Thaler v. Perlmutter,1 on March 18, 2025, supports the position adopted by the United States Copyright Office and is the latest chap…

Thaler v. Perlmutter is an a weird case because Thaler explicitly disclaimed human authorship and tried to register a machine as the author. Whereas someone trying to copyright LLM output would likely insist that there is human authorship is via the choice of prompts and careful selection of the best LLM output. I am not sure if claims like that have been tested.

The US copyright office has published a statement that they see AI output analogous to a human contracting the work out to a machine. The machine would hold the copyright, but can't, consequently there is none. Which is imho slightly surprising since your argument about choice of prompt and output seems analogous to the argument that lead to photographs being subject to copyright despite being made by a machine.

On the other hand in a way the opinion of the US copyright office doesn't matter, what matters is what the courts decide

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#466

Earlier quoted context omitted.

The argument is that converting static text into an LLM is sufficiently transformative to qualify for fair use, while distilling one LLM's output to create another LLM is not. Whether you buy that or not is up to you, but I think that's the fundamental difference.

> The U.S. Court of Appeals for the D.C. Circuit has affirmed a district court ruling that human authorship is a bedrock requirement to register a copyright, and that an artificial intelligence system cannot be deemed the author of a work for copyright purposes > The court’s decision in Thaler v. Perlmutter,1 on March 18, 2025, supports the position adopted by the United States Copyright Office and is the latest chap…

>I, like many others, believe the only way AI won't immediately get enshittified is by fighting tooth and nail for LLM output to never be copyrightable

If the person who prompted the AI tool to generate something isn't considered the author (and therefore doesn't deserve copyright), then does that mean they aren't liable for the output of the AI either?

Ie if the AI does something illegal, does the prompter get off scot-free?

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#467
post #456

Earlier quoted context omitted.

That sounds like wishful thinking. Every client I work for wants to reduce the rate at which humans need to intervene. You might not want that, but odds are your CEO does. And babysitting intermediate stages is not productive use of developer time.

And the odds are good you use the models and understand them in detail while the CEO is just buying the hype, ill informed or not.

Well, I want to reduce the rate at which I have to intervene in the work my agents do as well. I spend more time improving how long agents can work without my input than I spend writing actual code these days.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#468

GLM-4.7-Flash was the first local coding model that I felt was intelligent enough to be useful. It feels something like Claude 4.5 Haiku at a parameter size where other coding models are still getting into loops and making bewilderingly stupid tool calls. It also has very clear reasoning traces that feel like Claude, which does result in the ability to inspect its reasoning to figure out why it made certain decisions…

And you can run quantized versions on old hardware! Like 10 year old hardware. You might only get 3 tokens/sec, but it works.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#469
post #215

Pelican generated via OpenRouter: https://gist.github.com/simonw/cc4ca7815ae82562e89a9fdd99f07... Solid bird, not a great bicycle frame.

That's a bike that's ergonomically designed for pelicans.

It is unreasonable to expect pelicans to ride human bikes, they have different anatomy.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#470
post #29

Grey market fast-follow via distillation seems like an inevitable feature of the near to medium future. I've previously doubted that the N-1 or N-2 open weight models will ever be attractive to end users, especially power users. But it now seems that user preferences will be yet another saturated benchmark, that even the N-2 models will fully satisfy. Heck, even my own preferences may be getting saturated already. Op…

> No end-user on planet earth will suffer a single qualm at the notion that their bargain-basement Chinese AI provider 'stole' from American big tech.

Just like nobody cares[0] that American big tech stole from authors of millions of books.

[0] Interestingly, the only ones that cared were the FB employees told to pirate the Library Genesis and reporting back that "it didn't feel right".

Post reply on HN