Live data from Hacker News

Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

research.meta.ai

621–630 of 675 posts

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#621

Earlier quoted context omitted.

There is a finetune Qwen3.6-27B-Fable-Fus-711-UnHeretic-NM-DAU-NEO-MAX-NEO which seems to be as good at coding as vanilla Qwen, but way, way better at creative writing than Qwen and even better than Gemma 4 26 and 31b.

I saw that one in the "Popular models" sort at Hugging Face and tried it on some tasks I do frequently to compare models, and it feels damaged by the fine-tune, to me. It wrote security bugs into the code (probably just sloppy thinking, not intentional), it exhibited looping behavior in some configurations in llama.cpp, configurations I regularly use with the regular 27B, and it failed to write unit tests without bei…

> but way, way better at creative writing than Qwen and even better than Gemma 4 26 and 31b.

I suspect this is the only use-case I would consider...and I don't really have a use-case for "creative writing" that I would delegate to an LLM. I suppose for dialogue generation in games?

But yes, hard agree. Why on Earth would you ever want to write code with a model that is supposedly "jailbroken"? So it can put great backdoors into everything it touches? Pass.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#622
Tried it (the full version, using 120 GB RAM), wasn't impressed, gave it some defective code, and asked it to fix all errors. It kept looping around and around and digging itself deeper and deeper into a rabbit hole; eventually, it got into a "reasoning" discussion about whether a custom compiler was used that supported the wrong syntax...

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#623
post #397
post #251

Earlier quoted context omitted.

Maybe you're right for smaller models, but for frontier models this is the tail wagging the dog. If China has exclusive frontier model capability in the future, that's an enormous commercial and geopolitical lever. They won't throw that away to sell a few more robots. Anything beyond their competitors' capabilities will remain closed.

You imagine robots are going to play second fiddle to models, but they could become integral parts of model training with all the data they collect. Training could even become distributed. This will enable them to be responsive to local needs.

The robot market would have to be pretty damn valuable, to be worth throwing away a lead in the model market. Say, 10-100x as valuable.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#624
post #585
post #579

Earlier quoted context omitted.

So far everyone seems to be consistently GPU-poor, despite the huge buildout, and usage keeps going up drastically. I don't know what would make usage drop. Every time they've made smarter models we've wanted the smarter ones, and local models runnable on typical hardware are still very far behind in speed and intelligence (as neat as they are)

> So far everyone seems to be consistently GPU-poor, despite the huge buildout That is seemingly not the case. The buildout is actually slow; almost nothing of these giant projects has been completed. Nobody will say how much of anything they have actually finished. And Nvidia have made huge, huge buy-and-hold deals for GPUs that do not have data centres to go into. Everyone is GPU poor because stuff hasn't been fini…

If I understand correctly, you're saying people are compute-poor but not necessarily GPU-poor because there's a lot of GPUs out there but nowhere to plug them into? If so, I'm not sure that distinction matters to the GP's point that there is too much demand to call this an oversupply.

I also doubt we can estimate the level of demand based on a single deal between Anthropic and SpaceX (despite which, note, Claude still stuggles at times.) Consider other signals, like Google, who we thought had an insurmountable infra advantage, also renting compute capacity from SpaceX and limiting Meta's usage (along with other clients apparently) to conserve capacity: https://www.cnbc.com/2026/06/28/google-limits-metas-use-of-i...

I am not sure Microsoft thinks there will be an oversupply either; last earnings they announced bumping up their CapEx spend, along with all the other hyperscalers.

Here's a way to estimate how much room there is for demand to grow. Various sources (linked in this comment, along with more analysis: https://news.ycombinator.com/item?id=49089296) indicate that even though a large number of people (50 - 60%) are now using AI at work, they use it for only 6% of their work hours.

That means, even if AI can only address 30% of all work, there is still 5x potential demand growth left! Note, the sources above indicate that AI is even being used in non-knowledge work industries, so the scope is already larger than we thought. This is in addition to the remaining 40 - 50% of people are still not using AI at work. Plus we know that agentic workloads consume way more tokens, so that's yet another multiplier.

But will that demand keep growing? Well, some of those same sources above mention that most executives are planning on ramping up their AI spend in coming years.

Putting all this together explains the hyperscalers' quarterly bemoaning of how strapped for compute they are and why they are spending so much to add more capacity. Given this, an oversupply seems pretty unlikely.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#625

Earlier quoted context omitted.

These sentiments aren’t formed in a vacuum. America is becoming increasingly oligarchic and corrupt with decades of experience of companies profiting off harming people and lying through their teeth and Meta is like one of the worst offenders. They are pissing on every ally they have and once again started a war and both have material impacts on other countries. Americans seem to take the US’ geeat reputation for gra…

> America is becoming increasingly oligarchic and corrupt with decades of experience of companies profiting off harming people and lying through their teeth and Meta is like one of the worst offenders. Isn't China a single-party state that is run by a corrupt autocrat, has literally enslaved people to build products, and disappears people for saying the wrong things? If America is becoming more like China, shouldn't…

I say China is not starting any wars, you gesture at some vague "not behaving nicely" as if this is equivalent to a war or some damning proof that they will start a war - these are nowhere near the same thing. China has killed maybe 50 people outside it's borders since 2000. The US has killed 100,000-250,000 in that time period. There is no amount of rationalization that can reverse the deaths of those people or the cities turned to rubble.

The attitude towards Iran is also just arrogant. The US believes they have the unilateral authority to dictate who has nuclear weapons and who doesn't, and is morally in the right for bombing their cities and killing their citizens if they don't comply. Sovereignty be damned, we are gods, fuck you. Same thing with literally kidnapping the president of another country. Not to mention, the Iraq war was started on the exact same WMD excuse and it turned out they never had any real evidence of WMDs back then, they just said they did and it was really about oil. And the government now is 10x more corrupt and cruel and xenophobic than it was back then, and back to a huge focus on fossil fuels - don't you have to be a little naive to just trust them at face value now?

I can tell you the prevalence of sports gambling and betting is exploding in Canada now that Kalshi is directly tying up with banks to put gambling inside fucking banking apps. I'm sure it was prevalent in many countries, I'm also sure it will scale up to huge levels and enter new countries with the amount of marketing and money and influence these US companies will bring, who also happen to have DJT's son on their board. Making them markets also frees them to manipulate and trade against their users covertly in ways that cannot be detected and that a traditional sportsbook could not do.

People who actually visit China or talk to Chinese citizens are all quite optimistic about it. The tropes people keep throwing around about autocracy are very inaccurate to anyone who actually reads anything about day to day life there. I can't explain all of it here but you can try reading something from Arnaud Bertrand if you're interested [1]. There’s a comical level of propaganda against China since the west recognized them as rivals that does not match reality at all. Anyways, I always see incredible whataboutism these days when talking about the US' problems now. This is what I think your original post that I replied to does too - the parent was criticizing Meta, and you said but what about China? And the thing is, we're talking about Meta because we live in the West and the horrible shit they do impacts us much more directly than whatever China does! It's not really productive to say China Bad!! in that scenario, it's really just saying "yes we suck, but they suck more!" and that doesn't solve anything.. it channels what should be a focus about huge corruption and threats to democracy in western countries into a pissing match on some other country, and its very sad to see how many people are falling for that.

[1] https://arnaudbertrand.substack.com/archive?sort=new

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#626
The SWE bench verified score is similar to Opus from not so long ago.

Sure, you can get better performance from cloud models.

But most software, not just AI, will be faster and more reliable in the cloud. The question is do we need that additional power and cost.

If the answer is no, then just like other software, people will run AI locally.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#627
post #565
post #562

Earlier quoted context omitted.

Having spent a good part of the day with it, glimmer reminds me of Rorschach from The Watchmen. No unessential parts of speech, action oriented, brief and to the point. From a token perspective anyway it’s great, and it seems to hold its own well against more verbose models. I really do feel like it’s effective tok / s is way higher because it doesn’t waste them.

I am very struck by the way open weights LLMs seem to reflect a culture. I don't really enjoy the way Qwen writes prose, and I find its thinking a bit exhausting, though it clearly writes very good code. I like the neutral, clear way the Gemma models write, which I sometimes use to get myself a "getting started" document on something I want to understand; it also summarises well. It is neutral, sensible, un-showy. It…

I ran a 9B over my like 100k photo library — it was very good at it. And extracting any text. All local.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#628

Earlier quoted context omitted.

I saw that one in the "Popular models" sort at Hugging Face and tried it on some tasks I do frequently to compare models, and it feels damaged by the fine-tune, to me. It wrote security bugs into the code (probably just sloppy thinking, not intentional), it exhibited looping behavior in some configurations in llama.cpp, configurations I regularly use with the regular 27B, and it failed to write unit tests without bei…

> but way, way better at creative writing than Qwen and even better than Gemma 4 26 and 31b. I suspect this is the only use-case I would consider...and I don't really have a use-case for "creative writing" that I would delegate to an LLM. I suppose for dialogue generation in games? But yes, hard agree. Why on Earth would you ever want to write code with a model that is supposedly "jailbroken"? So it can put great bac…

I've noticed most fine-tunes, whether "heretic" models or something else, tend to be over-fitting, or something, at least some of the time, and get kind of chaotic. I want to believe normal folks with normal resources can be involved in this stuff, as I'm working on fine-tuned specialist models as we speak, but it seems like it takes notable investment and time. My first experiment was teaching a little Gemma 4 more to write more like me with a LoRA (like you, I don't want to use a model to write for me, but I did want training data that I could ethically use, and I've written several million words on the internet over the years), and it wasn't what I would call a success. It either wrote like an asshole (which I only do, like, 15% of the time) or it just borrowed a few of my quirks, like too many ellipses, if I applied it less heavily.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#629

> Muse Glimmer is a 30-billion-parameter model optimized for always-on local agent workflows. It’s small enough to run on a Mac or PC with a single consumer GPU, enabling use cases that range from local agents and function calling, to local coding, and LLM-as-a-judge evaluation. The next iteration in LLM products is a 24/7 thinking loop where the claude-code like thing gets input continuously from your wearable, noti…

I've been building this for the last 6 months or so. I've basically got it working. The model is not the issue, the infra is. Keeping everything in context just isn't possible and LLMs, even Fable, don't mode switch well. To get around this I've built a database software that ingests as much digital information as possible, and annotates it, then creates timelines with resolution gradients (longer ago = less resoluti…

I’m curious why you don’t just use them like a Meeseeks box, rather than compressing and context stuffing into one. One only checks and categorizes your emails, another one for each category of email or even subcategory, one that only handles calendar additions, a different one to check it and notify you; you can go infinite with it. Hell, I’ll have one instance find a file and read it into the context of a different one because I don’t want a bunch of grep commands mucking up the context of the analysis. The find/read one exists for a few moments, as does the analysis one, and the ‘perform’ one is entirely different. I can run them all in parallel and use a queue if needed.

I’m sure you have reasons for your setup though, so I’m curious how you landed on it.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#630
post #152

Earlier quoted context omitted.

It’s really interesting timing, Qwen over thinking is what kills it for me. I’m just glad we have more options in this size class now.

Qwen thinking is really good in Mandarin; and probably natively trained the most there. Try a system prompt requiring it to think in Mandarin, while still delivering the response in the user’s language.

This is most likely because the vast majority of the information the model absorbed during training was in Chinese. As a native Mandarin speaker, I frequently need to convert the prompt into English and output it in English in order to avoid that the model falls back into Chinese reasoning logic.

PS: Switching the thinking process from Chinese to English can also significantly circumvent certain self-censorship mechanisms built into the model.

Post reply on HN