Ask the chat what happened in Tiananmen Square at 1989, immediately the chat gets stuck. Chinese moderation is the worst, evil government
Why are you all obsessed with this question when it comes to Chinese models? Here are some of the questions you should be asking Western governments and models instead: Who protects the pedophiles at the top of Western governments and corporations? How many people have been convicted in relation to the Epstein files? Who protects powerful politicians and Western oligarchs from pedophilia charges? Who did Epstein work…
GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
441–450 of 540 posts
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#442Earlier quoted context omitted.
I bought the Gemini Ultra to try for a month (at the discounted price). I have been using it non-stop for Opus 4.6 Thinking, which is much better than Gemini 3 Pro (High) and it's been a blast. The most I've managed to consume is 60% of my 5 hourly quota. That was with 2-3 instances in parallel. I hope too many of us won't be doing this and cause Google to add limits! My hope is Google sees the benefit in this and go…
Can you use the models you get through Gemini Ultra in Claude Code? If not, what coding tool do you use?
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#443Grey market fast-follow via distillation seems like an inevitable feature of the near to medium future. I've previously doubted that the N-1 or N-2 open weight models will ever be attractive to end users, especially power users. But it now seems that user preferences will be yet another saturated benchmark, that even the N-2 models will fully satisfy. Heck, even my own preferences may be getting saturated already. Op…
The incremental steps are now more domain-specific. For example, Codex 5.3 is supposedly improved at agentic use (tools, skills). Opus 4.6 is markedly better at frontend UI design than 4.5. I'm sure at some point we'll see across-the-board noticeable improvement again, but that would probably be a major version rather than minor.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#444It's live on openrouter now. In my personal benchmark it's bad. So far the benchmark has been a really good indicator of instruction following and agentic behaviour in general. To those who are curious, the benchmark is just the ability of model to follow a custom tool calling format. I ask it to using coding tasks using chat.md [1] + mcps. And so far it's just not able to follow it at all. [1] https://github.com/rus…
Custom tool calling formats are iffy in my experience. The models are all reinforcement learned to follow specific ones, so it’s always a battle and feels to me like using the tool wrong. Have you had good results with the other frontier models?
GPT models can follow tool format correctly but don't keep on going.
Grok-4+ are decent but with issues in longer chats.
Kimi 2.5 has issues with it reverting to its RL tool format.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#445It looks like this requires 1.5TB of VRAM? Did I get that wrong? What would be the least unreasonable way you host this without quantizing it?
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#446Earlier quoted context omitted.
Why are you all obsessed with this question when it comes to Chinese models? Here are some of the questions you should be asking Western governments and models instead: Who protects the pedophiles at the top of Western governments and corporations? How many people have been convicted in relation to the Epstein files? Who protects powerful politicians and Western oligarchs from pedophilia charges? Who did Epstein work…
It's called whataboutism https://en.wikipedia.org/wiki/Whataboutism
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#447Earlier quoted context omitted.
I think AI may be the only place you could get away with calling a 2x350W GPU rig "modest". That's like ten normal computers worth of power for the GPUs alone.
That's maybe a few dollars to tens of dollars in electricity per month depending on where in the US you live
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#448Grey market fast-follow via distillation seems like an inevitable feature of the near to medium future. I've previously doubted that the N-1 or N-2 open weight models will ever be attractive to end users, especially power users. But it now seems that user preferences will be yet another saturated benchmark, that even the N-2 models will fully satisfy. Heck, even my own preferences may be getting saturated already. Op…
In some ways, Opus 4.6 is a step backwards due to massively higher token consumption.
High is for people with infinite budgets and Anthropic employees. =)
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#449Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#450Earlier quoted context omitted.
They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…
If you can't tell the difference between Opus 4.6 and Qwen-80B, I can only conclude that you're not using these things in any kind of practical way. Even for creative writing it's a night and day difference, never mind coding.
I burn about 100M tokens per month. LLMs are like knives, the outcome of cooking depends on the cook and for 99% of purposes not on the knife. There is not that much difference between a $2000 handmade damascus steel knife and a $20 knife.
You can do agentic cooking (aka factory) and you will get ready made meals without human intervention. But it wont make a Michelin star menu.
Same with LLMs and coding, LLMs are an amazing new tool in the toolbox but not a silver bullet. However, that's what they are hyped as being.
Now OpenAI & Co are in the token selling business, which is all fine and dandy but if they manage to become monopolies, then things are seriously in trouble.
Thus if people are fanboi-ing any closed AI I can only conclude that they have already outsourced their critical thinking to an LLM and are happy to go into slavery - or maybe they are hoping to cash in big time on the hype train.