Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

431–440 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#431
post #97

Earlier quoted context omitted.

In Silicon Valley we pay PG&E close to 50 cents per kWh. An RTX 6000 PC uses about 1 kW at full load, and renting such a machine from vast.ai costs 60 cents/hour as of this morning. It's very hard for heavy-load local AI to make sense here.

Yikes.. I pay ~7¢ per kWh in Quebec. In the winter the inference rig doubles as a space heater for the office, I don't feel bad about running local energy-wise.

God bless Canada. I love our cheap hydro power. <3

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#433
post #54

Earlier quoted context omitted.

I have no idea how an LLM company can make any argument that their use of content to train the models is allowed that doesn't equally apply to the distillers using an LLM output. "The distilled LLM isn't stealing the content from the 'parent' LLM, it is learning from the content just as a human would, surely that can't be illegal!"...

When you buy, or pirate, a book, you didn't enter into a business relationship with the author specifically forbidding you from using the text to train models. When you get tokens from one of these providers, you sort of did. I think it's a pretty weak distinction and by separating the concerns, having a company that collects a corpus and then "illegally" sells it for training, you can pretty much exactly reproduce t…

Contracts can't exclude things that weren't invented when the contracts were written.

Ultimately it's up to legislation to formalize rules, ideally based on principles of fairness. Is it fair in non-legalistic sense for all old books to be trainable-on, but not LLM outputs?

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#434

Earlier quoted context omitted.

I spent $10 in 2 minutes with that and gave up

Their 50 USD per month plan gives you 24M tokens per day: https://www.cerebras.ai/pricing

I had that for a few months and cancelled. They have minutely rate limits as well so you get 3-4 hyperspeed responses and then a 45 second pause waiting for the throttling to let your next request through.

And then, depending on what you're working on, the 24M daily allotment is gone in under an hour. I regularly burned it in about 25 minutes of agent use.

I imagine if I had infinite budget to pay regular API rates on a high usage tier, it would be really quite good though.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#435
post #20

Earlier quoted context omitted.

sleeper agents to do what? let's see how far you can take the absurd threat porn fantasy. I hope it was hyperbole.

There was research last year [0] finding significant security issues with the Chinese-made Unitree robots, apparently being pre-configured to make it easy to exfiltrate data via wi-fi or BLE. I know it's not the same situation, but at this stage, I wouldn't blame anyone for "absurd threat porn fantasy" - the threats are real, and present-day agentic AI is getting really good at autonomously exploiting vulnerabilities…

I could say that about Cisco and I would not be wrong.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#436

Earlier quoted context omitted.

Be careful with openrouter. They routinely host quantized versions of models via their listed providers and the models just suck because of that. Use the original providers only.

I specifically do not use the CN/SG based original provider simply because I don't want my personal data traveling across the pacific. I try to only stay on US providers. Openrouter shows you what the quantization of each provider is, so you can choose a domestic one that's FP8 if you want

Funny, living in Europe, I prefer using EU and Chinese hosts because as I don't want my data going to the US.

The trust in US firms and state is completely gone.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#437

Earlier quoted context omitted.

Anthropic has very tight limits, so you're basically using the worst (pricing-wise) SOTA cloud model as your baseline. I have $200 subs for both Claude and OpenAI, and I also bump into limits with Claude all the time, whether coding or research. With Codex, I ran into the limit once so far, and that's in a month of very heavy (sometimes literally 24 hours around the clock, leaving long-running tasks overnight) use.

I bought the Gemini Ultra to try for a month (at the discounted price). I have been using it non-stop for Opus 4.6 Thinking, which is much better than Gemini 3 Pro (High) and it's been a blast. The most I've managed to consume is 60% of my 5 hourly quota. That was with 2-3 instances in parallel. I hope too many of us won't be doing this and cause Google to add limits! My hope is Google sees the benefit in this and go…

Can you use the models you get through Gemini Ultra in Claude Code? If not, what coding tool do you use?

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#438
post #174

Really impressive benchmarks. It was commonly stated that open source models were lagging 6 months behind state of the art, but they are likely even closer now.

LLM benchmarks are largely irrelevant when it comes to "state of the art". They tell you if the model does poorly, but they are not at all a reliable signal of whether it does well. Open-weights models are still lagging quite a bit behind SOTA. E.g. there's still no open model that can match GPT-5 Pro or Gemini 2.5 Pro, and the latter is almost a year old by now.

Not true. For example, I think Gemini 3 Pro also can't match Gemini 2.5 Pro. Without benchmarks, it's just personal taste.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#440

Ask the chat what happened in Tiananmen Square at 1989, immediately the chat gets stuck. Chinese moderation is the worst, evil government

Why are you all obsessed with this question when it comes to Chinese models? Here are some of the questions you should be asking Western governments and models instead: Who protects the pedophiles at the top of Western governments and corporations? How many people have been convicted in relation to the Epstein files? Who protects powerful politicians and Western oligarchs from pedophilia charges? Who did Epstein work for, and why (hint: it’s not Russia or China)?
Post reply on HN