Earlier quoted context omitted.
Even if you had a 64GB machine: Are you willing to reserve 90% of your memory to run a LLM? With dirt cheap models like deepseek-v4-flash that will run "forever" on $10, the answer for me is clearly: no.
"With dirt cheap models like deepseek-v4-flash that will run "forever" on $10, the answer for me is clearly: no." When it's free, you are the product.
Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
31–40 of 681 posts
Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#32Will be interesting to see how Qwen3.8 27B compares against this once it releases this week. Seems like dense 30B is back in fashion? EDIT: An open weight version of Muse Spark 1.2 is going to be released as well: https://x.com/alexandr_wang/status/2086756152034066792 https://xcancel.com/alexandr_wang/status/2086756152034066792
UPD. was wrong on smaller, it's actually much larger
Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#33Earlier quoted context omitted.
Even if you had a 64GB machine: Are you willing to reserve 90% of your memory to run a LLM? With dirt cheap models like deepseek-v4-flash that will run "forever" on $10, the answer for me is clearly: no.
"With dirt cheap models like deepseek-v4-flash that will run "forever" on $10, the answer for me is clearly: no." When it's free, you are the product.
Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#34Will be interesting to see how Qwen3.8 27B compares against this once it releases this week. Seems like dense 30B is back in fashion? EDIT: An open weight version of Muse Spark 1.2 is going to be released as well: https://x.com/alexandr_wang/status/2086756152034066792 https://xcancel.com/alexandr_wang/status/2086756152034066792
Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#35Still needs 32-64GB memory to run it locally. 64GB Macbook pro with an M5 chip costs more than 4k Euros in Germany. A more practical model would be a language specific (e.g Python or JVM language) and excellent at tool calling and reasoning. Maybe that way they can shrink it even more.
Even if you had a 64GB machine: Are you willing to reserve 90% of your memory to run a LLM? With dirt cheap models like deepseek-v4-flash that will run "forever" on $10, the answer for me is clearly: no.
Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#36Still needs 32-64GB memory to run it locally. 64GB Macbook pro with an M5 chip costs more than 4k Euros in Germany. A more practical model would be a language specific (e.g Python or JVM language) and excellent at tool calling and reasoning. Maybe that way they can shrink it even more.
I'm sooo happy I pulled the trigger on upgrading and getting a new laptop (with 64 GB RAM) last summer. Feels like it was just in time before the exponential price jumps.
Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#37Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#38The favourable comparisons to Gemma 4 and qwen3.6 look promising!
Dense model makes it dog slow on anything without HBM. Max 15tok/sec on decode on DDR5 systems like a Spark or a Strix Halo -- and that's at 4 bit quant.
Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#39Still needs 32-64GB memory to run it locally. 64GB Macbook pro with an M5 chip costs more than 4k Euros in Germany. A more practical model would be a language specific (e.g Python or JVM language) and excellent at tool calling and reasoning. Maybe that way they can shrink it even more.
> We quantize weights to ~4-bit, bringing the LM under 20 GB. We validated minimal to no degradation on agentic tasks under compression.
https://www.reddit.com/r/LocalLLaMA/comments/1vkgsum/introdu...