Earlier quoted context omitted.
In Silicon Valley we pay PG&E close to 50 cents per kWh. An RTX 6000 PC uses about 1 kW at full load, and renting such a machine from vast.ai costs 60 cents/hour as of this morning. It's very hard for heavy-load local AI to make sense here.
Yikes.. I pay ~7¢ per kWh in Quebec. In the winter the inference rig doubles as a space heater for the office, I don't feel bad about running local energy-wise.
GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
431–440 of 540 posts
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#432Pelican generated via OpenRouter: https://gist.github.com/simonw/cc4ca7815ae82562e89a9fdd99f07... Solid bird, not a great bicycle frame.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#433Earlier quoted context omitted.
I have no idea how an LLM company can make any argument that their use of content to train the models is allowed that doesn't equally apply to the distillers using an LLM output. "The distilled LLM isn't stealing the content from the 'parent' LLM, it is learning from the content just as a human would, surely that can't be illegal!"...
When you buy, or pirate, a book, you didn't enter into a business relationship with the author specifically forbidding you from using the text to train models. When you get tokens from one of these providers, you sort of did. I think it's a pretty weak distinction and by separating the concerns, having a company that collects a corpus and then "illegally" sells it for training, you can pretty much exactly reproduce t…
Ultimately it's up to legislation to formalize rules, ideally based on principles of fairness. Is it fair in non-legalistic sense for all old books to be trainable-on, but not LLM outputs?
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#434Earlier quoted context omitted.
I spent $10 in 2 minutes with that and gave up
Their 50 USD per month plan gives you 24M tokens per day: https://www.cerebras.ai/pricing
And then, depending on what you're working on, the 24M daily allotment is gone in under an hour. I regularly burned it in about 25 minutes of agent use.
I imagine if I had infinite budget to pay regular API rates on a high usage tier, it would be really quite good though.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#435Earlier quoted context omitted.
sleeper agents to do what? let's see how far you can take the absurd threat porn fantasy. I hope it was hyperbole.
There was research last year [0] finding significant security issues with the Chinese-made Unitree robots, apparently being pre-configured to make it easy to exfiltrate data via wi-fi or BLE. I know it's not the same situation, but at this stage, I wouldn't blame anyone for "absurd threat porn fantasy" - the threats are real, and present-day agentic AI is getting really good at autonomously exploiting vulnerabilities…
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#436Earlier quoted context omitted.
Be careful with openrouter. They routinely host quantized versions of models via their listed providers and the models just suck because of that. Use the original providers only.
I specifically do not use the CN/SG based original provider simply because I don't want my personal data traveling across the pacific. I try to only stay on US providers. Openrouter shows you what the quantization of each provider is, so you can choose a domestic one that's FP8 if you want
The trust in US firms and state is completely gone.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#437Earlier quoted context omitted.
Anthropic has very tight limits, so you're basically using the worst (pricing-wise) SOTA cloud model as your baseline. I have $200 subs for both Claude and OpenAI, and I also bump into limits with Claude all the time, whether coding or research. With Codex, I ran into the limit once so far, and that's in a month of very heavy (sometimes literally 24 hours around the clock, leaving long-running tasks overnight) use.
I bought the Gemini Ultra to try for a month (at the discounted price). I have been using it non-stop for Opus 4.6 Thinking, which is much better than Gemini 3 Pro (High) and it's been a blast. The most I've managed to consume is 60% of my 5 hourly quota. That was with 2-3 instances in parallel. I hope too many of us won't be doing this and cause Google to add limits! My hope is Google sees the benefit in this and go…
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#438Really impressive benchmarks. It was commonly stated that open source models were lagging 6 months behind state of the art, but they are likely even closer now.
LLM benchmarks are largely irrelevant when it comes to "state of the art". They tell you if the model does poorly, but they are not at all a reliable signal of whether it does well. Open-weights models are still lagging quite a bit behind SOTA. E.g. there's still no open model that can match GPT-5 Pro or Gemini 2.5 Pro, and the latter is almost a year old by now.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#439Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#440Ask the chat what happened in Tiananmen Square at 1989, immediately the chat gets stuck. Chinese moderation is the worst, evil government