Live data from Hacker News

Muse Spark: Scaling towards personal superintelligence

ai.meta.com

301–310 of 392 posts

Re: Muse Spark: Scaling towards personal superintelligence

#301
post #266
post #166

Earlier quoted context omitted.

That fab will never be delivered. In five years you might see the manufacturing equivalent of a person dancing in spandex.

The only company that musk own and actually achieve something is spacex. so I believe you. He likes to hype things beyond what is actually possible. spacex is engineering masterpiece with how they revolutionize the space industry.

Do you know about his other company and what they do?

Re: Muse Spark: Scaling towards personal superintelligence

#302
post #279

The real question for me, if we assume they once again have a competitive frontier model, is what this means for Meta's strategy now. In particular, have they abandoned all their philosophy of the open ecosystem / open model play they were pursuing before? While it's true, llama4 sucked, I still can't help feeling they have lost ground compared to where they would have been if they maintained that strategy. Due to ll…

no need for an open harness when anthropic so kindly gifted the community theirs :)

Re: Muse Spark: Scaling towards personal superintelligence

#303
post #46

Ran some of my internal benchmarks against this and I'm very unimpressed. I don't think this moves them into the OAI v Anthropic v Gemini conversation at all. Major analytical errors in their response to multiple of my technical questions.

even gemini is not in that conversation

Re: Muse Spark: Scaling towards personal superintelligence

#304
post #203

Earlier quoted context omitted.

I suspect this is the real reason behind Anthropic limiting subscriptions to their own products and keeping API prices several times higher than comparable models. Applications more sticky than API users and less technical users more sticky than programmers (ie Cowork more sticky than Code).

Anthropic generally seem more into living within market discipline and market signals of some sort. Products with margins, even if it's sort of irrelevant considering R&D costs and capital inflow. That said, there's nothing like the real thing. The risk is something like the railroad bubble and the dotcom. Over-investement, circular revenue and a timeline that doesn't work. Or, maybe it'll work out.

Maybe they’ll figure out how to make an agent train an agent.

Re: Muse Spark: Scaling towards personal superintelligence

#306

Earlier quoted context omitted.

Alexandr Wang suggesting this might be open-weights/source in the future gives me hope. Hopefully they stay on this path.

I have a feeling it won't be this exact model, but rather smaller distilled variants, similar to the gemma line

It is fair to think so because that is what everyone is doing. But being Meta and considering Llama, if MSL is going to keep releasing models and wants to join back the AI war, they may actually open weights just to get more attention. Once they establish a sizable community, they can start guarding their frontier models.

Re: Muse Spark: Scaling towards personal superintelligence

#307
post #287
post #239

Earlier quoted context omitted.

Post actual results, make a blog post. Don't just say "this sucks" without tangible evidence. Otherwise you're doomed to "sample size of one" level of relevance.

I may already have but I'm pseudonymous on this website.

[flagged]

Re: Muse Spark: Scaling towards personal superintelligence

#308
post #92

I don't get the comments trashing this. If it slightly beats or even matches Opus 4.6, it means Meta is capable of building a model competitive with the leading AI company. Sure, they spent a lot of money and will have on-going costs. But how much more work would it take to turn that into a coding agent people are willing to try (and pay for) along side their usage of a collection of agents (Claude, Codex, etc)? Also…

Comments trashing this are rightly correct skeptics who remember the benchmaxxing of llama 4. This model was out in the woods as early as like a couple months ago but they didn't release it because it was at gemini 2.5 pro levels.

> 4. This model was out in the woods as early as like a couple months ago but they didn't release it because it was at gemini 2.5 pro levels.

Source? (Even if rumor)

Re: Muse Spark: Scaling towards personal superintelligence

#309
post #92

I don't get the comments trashing this. If it slightly beats or even matches Opus 4.6, it means Meta is capable of building a model competitive with the leading AI company. Sure, they spent a lot of money and will have on-going costs. But how much more work would it take to turn that into a coding agent people are willing to try (and pay for) along side their usage of a collection of agents (Claude, Codex, etc)? Also…

Why go into coding agents? Both anthropic and OpenAI are going all in on that. The opportunity is customer facing AI now. OpenAI has the mindshare but they going to have to decide if they allocate their limited compute for free users or go all in trying to keep up with Anthropic in enterprise.

If you squint at coding agents you see the next OS.

Maybe better phrasing is “HCI paradigm”, but that somehow manages to say everything and nothing.

Re: Muse Spark: Scaling towards personal superintelligence

#310

Earlier quoted context omitted.

the models were objectively horrible

They really weren't horrible. They were ~gpt4o, with the added benefit that you could run them on premise. Just "regular" models, non "thinking". Inefficient architecture (number of active out of total) but otherwise "decent" models. They got trashed online by bots and chinese shills (I was online that weekend when it happened, it's something to behold). Just because they were non-thinking when thinking was clearly t…

> They were ~gpt4o, with the added benefit that you could run them on premise.

No, they are bad models. They were benchmaxxed on LMAreana and a few other benchmarks but as soon as you try them yourself they fall to pieces.

I have my own agentic benchmark[1] I use to compare models.

Llama-4-scout-17b-16e scores 14/25, while llama-4-maverick-17b-128e scores 12/25.

By comparison gemma-4-E4B-it-GGUF:Q4_K_M scores 15/25 (that is a 4B parameter model!) - even GPT3.5 scores 13/25 (with some adjustment because it doesn't do tool calling).

Llama 4 was a bad model, unfortunately.

[1] https://sql-benchmark.nicklothian.com/#all-data

Post reply on HN