Earlier quoted context omitted.
But they don't have the same human in the loop though.
that software is called autonomous agents, the term autonomous has nothing to do with human in the loop, it is the complete opposite.
GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
461–470 of 540 posts
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#462GLM-4.7-Flash was the first local coding model that I felt was intelligent enough to be useful. It feels something like Claude 4.5 Haiku at a parameter size where other coding models are still getting into loops and making bewilderingly stupid tool calls. It also has very clear reasoning traces that feel like Claude, which does result in the ability to inspect its reasoning to figure out why it made certain decisions…
minimax-m.2 is close
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#463Earlier quoted context omitted.
They had plenty of time to update their system prompts so they don't be embarrassed. I noticed whenever such meme comes out, if you check immediately you can reproduce it yourself, but after a free hours it's already updated.
thats not how it works
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#464Earlier quoted context omitted.
If that's what they're tuning for, that's just not what I want. So I'm glad I switched off of Anthropic. What teams of programmers need, when AI tooling is thrown into the mix, is more interaction with the codebase, not less. To build reliable systems the humans involved need to know what was built and how . I'm not looking for full automation, I'm looking for intelligence and augmentation, and I'll give my money and…
That sounds like wishful thinking. Every client I work for wants to reduce the rate at which humans need to intervene. You might not want that, but odds are your CEO does. And babysitting intermediate stages is not productive use of developer time.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#465Earlier quoted context omitted.
> The U.S. Court of Appeals for the D.C. Circuit has affirmed a district court ruling that human authorship is a bedrock requirement to register a copyright, and that an artificial intelligence system cannot be deemed the author of a work for copyright purposes > The court’s decision in Thaler v. Perlmutter,1 on March 18, 2025, supports the position adopted by the United States Copyright Office and is the latest chap…
Thaler v. Perlmutter is an a weird case because Thaler explicitly disclaimed human authorship and tried to register a machine as the author. Whereas someone trying to copyright LLM output would likely insist that there is human authorship is via the choice of prompts and careful selection of the best LLM output. I am not sure if claims like that have been tested.
On the other hand in a way the opinion of the US copyright office doesn't matter, what matters is what the courts decide
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#466Earlier quoted context omitted.
The argument is that converting static text into an LLM is sufficiently transformative to qualify for fair use, while distilling one LLM's output to create another LLM is not. Whether you buy that or not is up to you, but I think that's the fundamental difference.
> The U.S. Court of Appeals for the D.C. Circuit has affirmed a district court ruling that human authorship is a bedrock requirement to register a copyright, and that an artificial intelligence system cannot be deemed the author of a work for copyright purposes > The court’s decision in Thaler v. Perlmutter,1 on March 18, 2025, supports the position adopted by the United States Copyright Office and is the latest chap…
If the person who prompted the AI tool to generate something isn't considered the author (and therefore doesn't deserve copyright), then does that mean they aren't liable for the output of the AI either?
Ie if the AI does something illegal, does the prompter get off scot-free?
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#467Earlier quoted context omitted.
That sounds like wishful thinking. Every client I work for wants to reduce the rate at which humans need to intervene. You might not want that, but odds are your CEO does. And babysitting intermediate stages is not productive use of developer time.
And the odds are good you use the models and understand them in detail while the CEO is just buying the hype, ill informed or not.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#468GLM-4.7-Flash was the first local coding model that I felt was intelligent enough to be useful. It feels something like Claude 4.5 Haiku at a parameter size where other coding models are still getting into loops and making bewilderingly stupid tool calls. It also has very clear reasoning traces that feel like Claude, which does result in the ability to inspect its reasoning to figure out why it made certain decisions…
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#469Pelican generated via OpenRouter: https://gist.github.com/simonw/cc4ca7815ae82562e89a9fdd99f07... Solid bird, not a great bicycle frame.
It is unreasonable to expect pelicans to ride human bikes, they have different anatomy.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#470Grey market fast-follow via distillation seems like an inevitable feature of the near to medium future. I've previously doubted that the N-1 or N-2 open weight models will ever be attractive to end users, especially power users. But it now seems that user preferences will be yet another saturated benchmark, that even the N-2 models will fully satisfy. Heck, even my own preferences may be getting saturated already. Op…
Just like nobody cares[0] that American big tech stole from authors of millions of books.
[0] Interestingly, the only ones that cared were the FB employees told to pirate the Library Genesis and reporting back that "it didn't feel right".