Earlier quoted context omitted.
The researchers who published Attention Is All You Need didn’t have the benefit of the LLMs they birthed. Take a look at the prompt that solved the Cycle Double Cover conjecture, and which has been adapted to achieve breakthroughs in cybersecurity. The field is entering a feedback loop that is leading to exponential innovation. We’re at the beginning of the curve. And right now the big iron data center approach is br…
Because I wasn't familiar with it and others might be curious too: that prompt is available at https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98... and is just below 5 KiB of text.
Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
611–620 of 681 posts
Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#612even 30B model is too large to large on local device (low end). meta should provide free hosted model api to use it.
Meanwhile those of us with 128GB RAM plus some VRAM don't have any good modern (last 8 months) open weights models to make use of all that. I don't care if it would run 5 tok/s, I want a smarter model than Qwen3.6 which avoids loops and can handle more context than 80k before crashing.
Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#613Remember when we needed 200 servers for an enterprise website because Apache used one process or thread per connection - and Nginx collapsed that into a single box overnight? That moment for LLMs is near. It’s going to move us from the big iron era of AI to small portable brains. Nature has already proved it’s possible with 20 watts and very little heat generation. And I think the data center buildout will end in car…
Everyone keeps repeating this who doesn’t understand the underlying technology. Small llms are still way more efficiently server on big GPUs. Sharing server capacity takes advantage of the massive parallel throughput and sharing of memory bandwidth. You are sharing the GPUs with thousands of concurrent users.
Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#614Remember when we needed 200 servers for an enterprise website because Apache used one process or thread per connection - and Nginx collapsed that into a single box overnight? That moment for LLMs is near. It’s going to move us from the big iron era of AI to small portable brains. Nature has already proved it’s possible with 20 watts and very little heat generation. And I think the data center buildout will end in car…
Agree. I know enough about the vagaries of scientific progress to not put any money on any timeline but directionally, that's where we are headed.
> And I think the data center buildout will end in carnage.
Disagree. And this is quite the leap from the previous statement, btw. The carnage happens if the demand for general purpose GPU compute disappears and even then there are so many ways to salvage the asset.
Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#615Unsloth has quantized versions uploaded: https://huggingface.co/unsloth/Muse-Glimmer-30B-GGUF The quantized releases often change in the weeks following release as new improvements are discovered, so either use a tool that checks HuggingFace for new versions or manually check back in a few days or weeks to check for improved versions. Initial reports are good. It hasn't been out long enough for anyone to really test…
Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#616> Muse Glimmer is a 30-billion-parameter model optimized for always-on local agent workflows. It’s small enough to run on a Mac or PC with a single consumer GPU, enabling use cases that range from local agents and function calling, to local coding, and LLM-as-a-judge evaluation. The next iteration in LLM products is a 24/7 thinking loop where the claude-code like thing gets input continuously from your wearable, noti…
Help convince Firefox of this: https://news.ycombinator.com/item?id=46294238 Rather than develop its own AI, Firefox should develop a system to pipe your html rendered browsing history in real time so external local services can process it: https://connect.mozilla.org/t5/ideas/archive-your-browser-hi... . Firefox could be the only browser that does this.
Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#617Still needs 32-64GB memory to run it locally. 64GB Macbook pro with an M5 chip costs more than 4k Euros in Germany. A more practical model would be a language specific (e.g Python or JVM language) and excellent at tool calling and reasoning. Maybe that way they can shrink it even more.
I am running it on a single RTX 3090 (24GB VRAM). Some folks on Reddit are having the same experience: https://www.reddit.com/r/LocalLLaMA/comments/1vkm42m/muse_gl... It uses an order of magnitude less VRAM at longer contexts which is a huge advantage over Qwen 3.6 27B
Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#618Earlier quoted context omitted.
Yeah, but if it's huge, how many can run it? Many folks struggle to run 200B+ models
Two DGX sparks will run the full Deepseek V4 Flash full. (It is definitely expensive, but relatively easy and compact; and extremely power efficient)
Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#619Earlier quoted context omitted.
There is a big astroturfing going on social media platforms by the chinese. Did you notice 'day in a life of unmarried 30 yr old lady in china' videos flooding usa social media. Regular ppl in the west now hold mildly positive views of the ccp and how 'advanced' china is than usa. Then there are europeans who now are looking for china to give them the technology handout now that relationship with usa has soured.
The interesting thing is that China began as an agrarian economy and quickly modernized via influence/direction of the CCP. The United States used to have a dominant middle class that was geographically distributed (cities & rural areas inclusive), but the advent of the tech economy has also been having a similar effect here as it did in China: massive wealth accumulation in Tier 1-3 cities and everyone else being la…
Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
#620Earlier quoted context omitted.
I would hope that Qwen 3.8 is better. It's been 4 months, and we've seen almost no progress in this space. As people have called out, Glimmer appears to be a trade-off rather than a clear winner. And from what I've been reading, no one is expecting Qwen 3.8's model in this space to be a clear winner, but just slightly and marginally better. That's a little concerning as DeepSeek v4 Flash proved at it larger sizes the…
I don’t think four months without a major breakthrough is cause to abandon all hope just yet. ;) The wild pace of LLM development is highly atypical, and we’re still in the ‘initial rush’ phase of development. For contrast, the Newcomen steam engine (widely considered the first commercially useful engine) was used for over 60 years before the next major improvements. Now, 300 years later, we’re still finding ways to…
Off-topic, but I stumbled upon the first Newcomen engine imported into Australia in a museum in Sydney and I was unexpectedly charmed (not an Engine Guy). It's large, but nothing like the awe of "mega-engineering", it's crude, but it clearly has such amazing utility (when compared to a reality without it) and it changed the world