Live data from Hacker News

Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

research.meta.ai

181–190 of 682 posts

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#181
post #166

Another candidate for the 7900XT (20GB VRAM) I got sitting around. I pulled latest llama.cpp (targeting vulkan during build) after seeing a muse PR merged a few hours ago, and unsloth/Muse-Glimmer-30B-GGUF:UD-Q4_K_XL runs on my 7900XT barely (and with no MTP). Sits at 19GB VRAM w/ 4 parallel 113k context slots, all layers on GPU, and at 700 tok/s prompt, and ~36 tok/s generation. Waiting on Q3 to download to check sp…

Q3 results: unsloth/Muse-Glimmer-30B-GGUF:UD-Q3_K_XL gets down to 15.6GB VRAM and full context (131k) on the 4 parallel slots. Prompt/generation speeds about the same. Overall feeling like a nicer-fitting Qwen 3.6 27B, but want to test out MTP generation speeds once I can.

edit: My favorite bit of reasoning I saw go by in my "generate me a beautiful code snippet" anecdote: 'Could give a snippet of beautiful code: the "hello world" in brainfuck? No.'

edit2: my first dflash speculative model! no mtp. I'm up to ~60 tok/s on empty context with `--spec-type draft-dflash`

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#183

Earlier quoted context omitted.

What is the other company that you would never work for?

He said OpenAI in the comment (if I read it correctly)

Ah, my bad. Thanks for pointing that out :)

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#184
post #132
post #123

Earlier quoted context omitted.

Considering how all the big players are playing fast [1] and loose [2] with limits, billing [3] and adding undisclosed changes that burn your tokens on autopilot [4], it can't happen soon enough. [1]: Limits may change without notice, including due to capacity constraints. - https://support.google.com/gemini/answer/16275805?sjid=14713... . [2]: "standard limits" are never defined - https://support.google.com/gemini/a…

Not to mention all the other ways they can screw you: - Middle of the day, servers busy? Swap to Sonnet while pretending it's still Opus. Many people won't notice, and nobody can prove anything if they suspect. - Middle of the night, server load is light? Put it into extra thinky mode so it burns more tokens to ramp up the bills. Flip the switch where it gets really pedantic about writing lots of extra test cases and…

It all sounds like having to rely on a dodgy housing contractor that wants to steal from you, take shortcuts AND choose the gold-plated options from their supplier friends, and will start doing this the minute you are not on site supervising. You don't do it yourself (because the contractor is faster and stronger than you in many ways) but you can't leave, so you're stuck on the worksite just watching them.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#185

Earlier quoted context omitted.

[flagged]

Think real hard about that. What does it mean if the only hacker chat group on the planet despises meta this much? Think.

You think this is the only hacker chat group on the planet?

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#186

Great to see Meta back, looks like really strong, local model, can't wait for llama.cpp support.

some support already merged, and I verified in a local build that it runs (cannot get MTP params working tho, about ~40 tok/s on my beefy 800GB/s 7900XT w/ 20GB VRAM). https://github.com/ggml-org/llama.cpp/pull/26841

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#187
That's a bit amusing - not that I have the hardware to run it, but officially it's not available in Hong Kong. Not that getting it would be much of a problem with a help of a VPN either, but I'll assume mainland China is also restricted. Certainly not a competition for Chinese open weight models... in China.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#188
post #136
post #47

It is interesting but it does look like a careful distillation of (Spark and) biggers open-weight models. The progress compared to Qwen3.6 27B is good, not that impressive, it's a 4 months old model. (kuto to them to compare to 27B dense and not 35B MoE, it's more fair to do so). It is very probable that Qwen3.8 27B will crush Glimmer-30B on most benchmarks.

Still great if they want to play in this space. Having competition for the 24-32GB VRAM target is only good for the end user.

Agreed, the trend in this consumer-accessible range is encouraging.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#189
post #97

With the business model for API based LLMs looking iffy at best it seems like we’re heading back to the “server under your desk” era of IT again.

With how expensive consumer hardware is and will continue getting (due to LLM demand), good luck getting a "server under your desk" for something less than an arm, leg, and first born. Until A100 prices are reliably under 1.70$ an hour, there is no GPU/AI bubble and Michael Burry doesn't know anything about GPUs.

There are lots of points in a spectrum of choices. DGX Sparks, Strix Halos, and the surviving Mac Studios can easily run these 30B class models, just not as fast. So maybe just the leg, but you can keep the arm and first born.

And super noteworthy is that a 27B model (Qwen 3.6 27B) from this year is a huge improvement over a 120B model (gpt-oss:120b) from last year. The goal posts are moving, but at some point "good enough" is good enough for the kind programming I like to do.

Re: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

#190

Earlier quoted context omitted.

Meta can never be redeemed, but it's still valid to admit that FB at one point had a very badass engineering culture. They're one of 2 companies I would absolutely never work for (weapons etc aside). FB's recruiters hounded me so often I requested that they blackball me. The day they became Meta, I learned this by checking my email to see that they started trying to reach out again. I once again requested that they b…

> FB at one point had a very badass engineering culture Perpetually kneecapped by one of the worst management cultures I've ever seen

Would you say those badass engineers were/are managed by Careless People?
Post reply on HN