Earlier quoted context omitted.
Cool! I'm thinking about a local set up. What's your usual tokens/second rate?
Not OP, but I’m running local models on a M1 Max as well with 64GB RAM. It varies by model, but I’m getting 50-60 t/s with Qwen 3.6 35B and Qwen 3 coder 30B. I’ve also used Qwen 3.8 27B but I get 10t/s on it. It’s useable in some use cases, but I rely mostly on my $20 Claude subscription.
GLM-5.3 is now open-weight
271–280 of 298 posts
Re: GLM-5.3 is now open-weight
#272Re: GLM-5.3 is now open-weight
#273Re: GLM-5.3 is now open-weight
#274Earlier quoted context omitted.
With roughly 2.7 million seconds per month, times 14 tokens per second, you are getting 38.5 million tokens a month at most. That’s less than 164USD worth of GLM5.3 tokens on the inference market. So that 6000 USD rig will take 3 years to break even - and only if it runs continuously. And this is being generous, as it’s not even taking quantisation into account.
I think if you steel-man what I'm saying, what you're saying falls apart. 14 tokens per second was rare. It only dropped that low in one scenario where he had it single shot an entire game (flappy bird clone) from scratch, with different assets, all self created, and so on. It ended up resulting in the LLM doing stuff like plotting out a some odd 100 item long to-do list, requerying it repeatedly, and so on. And it s…
I agree that there are many other reasons than cost alone.
Re: GLM-5.3 is now open-weight
#275Earlier quoted context omitted.
> We're nowhere near a Fable-class model IMO, but things are going to get interesting in this next year. I'm wondering of you could clarify your thoughts on this. I've had a hard time evaluating what Fable-class actually is capable of that sets them (or really it) apart from other models in a very significant way.
Have you tried it out? I doubt there is a general accepted definition. For me it is just more capable of deep reasoning/handling complexity. Still can mess up, still does not count as strong AI - but a level above Opus and co.
I have a hard time getting models like GLM 5.3 to not perform on my tasks but I might be biased.
Re: GLM-5.3 is now open-weight
#276Earlier quoted context omitted.
Seems like it's easier to ask forgiveness than permission. And, I wouldn't bet on this legislature ever doing anything that would disempower fossil energy or reduce their profits, even a little bit.
There are good reasons you aren't allowed to plug random power generators into the grid. You might be allowed to have your own ones not connected to the grid. Remember that graphics cards run off poorly regulated 12V, although you'd want to regulate it anyway because they're expensive to replace if I'm wrong.
Re: GLM-5.3 is now open-weight
#277Earlier quoted context omitted.
There are good reasons you aren't allowed to plug random power generators into the grid. You might be allowed to have your own ones not connected to the grid. Remember that graphics cards run off poorly regulated 12V, although you'd want to regulate it anyway because they're expensive to replace if I'm wrong.
Nobody mentioned plugging into the grid.
Re: GLM-5.3 is now open-weight
#278Earlier quoted context omitted.
With roughly 2.7 million seconds per month, times 14 tokens per second, you are getting 38.5 million tokens a month at most. That’s less than 164USD worth of GLM5.3 tokens on the inference market. So that 6000 USD rig will take 3 years to break even - and only if it runs continuously. And this is being generous, as it’s not even taking quantisation into account.
I think the “killer app” is doing inference without sending the data to China or the US. At home it’s overkill but imagine you are an EU consultancy with a lot of client data to work on, or a company/institution with a lot of sensitive data, buying the hardware to make sure the data stays private is a huge benefit. So is that you “own” the model. Its capabilities, price or access don’t change at someone else’s whim.
Needing an EU native option is really the one and only reasonably objection I've heard against using LLMs from the cloud, the rest is tin-foil hat level unless you're actually intending to meddle with the inference or fine tuning or something beyond just querying.
Re: GLM-5.3 is now open-weight
#279Earlier quoted context omitted.
Have you already forgotten the Fable drama that happened just two months ago?
Yes, I remember recurring extensions on plan inclusion then becoming permanent - that's my point.
Re: GLM-5.3 is now open-weight
#280Earlier quoted context omitted.
I have a Strix Halo and dual 32GB GPUs in my desktop, that sit idle right now, because the electricity to run them and to cool them in 110F weather Texas is currently experiencing pretty much nulls any savings I might see over getting better models from cloud providers. While I mostly use Claude or Codex with subscriptions for agentic work, for API use DeepSeek has usually been my go to, but now I guess it's GLM 5.3…
Solar panels are produced for less than $1 per watt btw. Shame they're illegal in Texas because they compete with the governing oil industry.
Homeowners can’t put solar panels on their roof to use the produced electricity?