Live data from Hacker News

Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

research.meta.ai

421–430 of 682 posts

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#421

Earlier quoted context omitted.

> The major AI labs are gross profitable when selling access to inference. Do you have a good source for this?

Well, I can run some models that are better than some of the weaker and cheaper Anthropic models locally, like Haiku 4.5, and solve tasks that would cost ~4500$ every day in tokens, so yeah, they are definitely extremely profitable on inference.

Are you also paying 8,000 employees [1] and funding massive infrastructure [2]?

[1] https://www.makerstations.io/openai-employee-statistics/

[2] https://www.wheresyoured.at/oai_docs/

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#422

Earlier quoted context omitted.

While composing a reply to a comment throwing tons of shade on American AI, I took some time to check out the commenter’s HN profile. Their comment history was about 50% such comments. Their submission history started with an article about how Russia was unfairly blamed for some hacking campaign. It’s entirely possible that this is not a foreign influence campaign. Perhaps there’s a group here that is simply anti-Ame…

There's literally a Chinese-funded influence campaign against American AI: https://openai.com/index/disrupting-malicious-uses-of-ai-dat...

lol you believe this bullshit? "likely PRC-origin cluster" what an awesome amount of proof corporations need to convince the gullible.

Why not "OpenAI used simplified chinese to create a fake prc-origin campaign and media buzz to convince the public they actually love data-centers and anti-datacenter sentiment is a psyop" there's an equal amount of proof provided for either scenario.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#423
post #345

Earlier quoted context omitted.

I don't really use the Qwen 3.6 27B though I do test the variants (Bonsai, ThinkingCap). I really like the 3.6 35B A3B for experiments, and it seems OK, but as you say, it spins round in thinking loops more than say the 26B Gemma 4 does. If Muse doesn't actually-wait itself as much it will be very interesting. I am just downloading it to run my small tests.

the “Actually… But wait!” style responses are so annoying, even Claude opus struggles with this so I’d be interested if meta has done something to cut down on that while still giving good responses

It's not perfect but it is very terse! Better than BottleCap managed to do with post-training Qwen in ThinkingCap.

I suspect it will help a lot with enabling preserve-reasoning, because the biggest apparent limitation of this model is the 128K context window.

Though the practical issue I am seeing on my M1 Max MBP is that performance suddenly drops off a cliff if I have DFlash enabled.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#424

Earlier quoted context omitted.

> On the other hand, one should not discount the value of HN as tastemaker and trendsetter. I would encourage dedicated readers here to aggressively and persistently discount the value of HN as a tastemaker and trendsetter. HN is actually a trailing indicator on tastes and trends, essentially by design. Things only make it to the front page if they get submitted and voted upward by a large number of people. That mean…

So - presumably you have another site in mind that is better. I'd be intrigued to know which one you would recommend.

Sounds like something a bot would say!

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#425

Earlier quoted context omitted.

It’s really interesting timing, Qwen over thinking is what kills it for me. I’m just glad we have more options in this size class now.

I've been using Qwen3.6 35B A3B, and with reasoning turned on, I'd say 2/3 (give or take) of the tokens for a response are thinking tokens. Which at 70+ tps locally, that isn't that awful. I run an 80k context across 4-10 "agents" for my solo TTRPG, where Qwen is the GM, each NPC at a location, the director, and the narrator. Each turn is about 45-60 seconds to generate all of the various responses. The GM and direct…

Are you running inference in parallel? 70 tps seems low for parallel execution.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#426
post #76

Earlier quoted context omitted.

4K bucks buys you around 180 months of with zero upfront cost.

Problem is that might go away or get nerfed.

If that happens you can still buy hardware later with almost certainly more (tok/s)/$ and better capabilities to run newer models more efficiently (remember native MXFP4?). Right now basically every generation of accelerator is adding new capabilities. These aren't yearly DirectX 9.0c-compatible GPU performance bumps.

As an individual, for average privacy needs (e.g. open source or at-home coding and automation), it's pretty much complete nonsense financially to self-host LLMs currently or select hardware now based on the capability to do so, and pay thousands of bucks extra.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#427

Earlier quoted context omitted.

I'm a 52 year old natural born US citizen whose ancestors have been here for generations and I'm currently anti-American. Why wouldn't I be? We've never been the shining beacon of light we would claim to be, but we're so fucking awful now.

it is strange that when people are pro-american, or pro-any-ingroup, nobody asks for justification; OTOH, make a mistake of speaking from rational viewpoint based on historical events, expect a lot of hate from reddit/HN/any-ingroup-forum. I applaud your openness to speak. I've traveled/interacted with many nationalities, and very rarely I come across someone who is open/rational enough to openly state their dislike…

> is strange that when people are pro-american, or pro-any-ingroup, nobody asks for justification

Being pro a group can be strictly positive sum - wanting to lift that group up, likely because you consider yourself part of it or on the same team. It's possible your intentions are bad, but they certainly don't have to be.

Being anti a specific group is inherently negative. Perhaps they deserve it, but that requires justification in a way simply being positive does not.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#428
post #352

Earlier quoted context omitted.

From my perspective it doesn't make sense to talk about the number of parameters. What matters is model size in bytes and its performance at that certain size. Meta actually relesed official 4 bit quants in 17GB, but I haven't seen any indication that training was quant-aware, so the quants are not going to have same performance. 3.6 27B has official FP8 quant that AFAIR was trained with quantization awareness. The b…

> but I haven't seen any indication that training was quant-aware readme on huggingface says they've benchmarked the quants -- for 17GB quant reported 1% avg loss across 15 benchmarks (sadly no breakdown). I assume that's strong enough signal for QAT. Not just first party quants, but they cared to monitor degradation.

> sadly no breakdown

That's exactly the point. We know short context knowledge stuff does not regress with quantization. But I expect agentic intelligence to suffer greatly.

If I were to pick one bench, I would like to compare quants on TerminalBench Hard. But then Glimmer already loses to 3.6 27B on it by a large margin.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#429

Remember when we needed 200 servers for an enterprise website because Apache used one process or thread per connection - and Nginx collapsed that into a single box overnight? That moment for LLMs is near. It’s going to move us from the big iron era of AI to small portable brains. Nature has already proved it’s possible with 20 watts and very little heat generation. And I think the data center buildout will end in car…

First point is plausable, moving from bigger models to smaller models. But the nature thing is a bit of an overstatement, yes our brains are very efficient but they are fundamentally different from LLMs so it doesn't really map.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#430

Still needs 32-64GB memory to run it locally. 64GB Macbook pro with an M5 chip costs more than 4k Euros in Germany. A more practical model would be a language specific (e.g Python or JVM language) and excellent at tool calling and reasoning. Maybe that way they can shrink it even more.

I am running it on a single RTX 3090 (24GB VRAM).

Some folks on Reddit are having the same experience: https://www.reddit.com/r/LocalLLaMA/comments/1vkm42m/muse_gl...

It uses an order of magnitude less VRAM at longer contexts which is a huge advantage over Qwen 3.6 27B

Post reply on HN