Live data from Hacker News

Muse Spark: Scaling towards personal superintelligence

ai.meta.com

231–240 of 392 posts

Re: Muse Spark: Scaling towards personal superintelligence

#231

So this is why Anthropic rushed the weirdest "pre-responsible-disclosure-totally-not-for-marketing" announcement yesterday? To make sure Spark doesn't steal their thunder? (Spark beats Opus 4.6 on some benchmarks...). Or did I become a bitter cynical old man.

Last i checked with friends at meta they are pretty deeply invested in using claude for coding etc. anthropic has nothing to be scared of at MSL.

If spark beats opus 4.6, why is meta wasting money on opus internally?

Re: Muse Spark: Scaling towards personal superintelligence

#232

Personal Superintelligence made me think this was an open-source model being released and I was excited. Then I continued reading and I'll just wait until the model comes out.

I wonder if Zuck will ever internalize that the words ‘personal’ and ‘meta’ will not be taken seriously together for another decade (if they don’t make another gaff).

Re: Muse Spark: Scaling towards personal superintelligence

#233
post #11

I'm cautiously waiting for the feedback from the first users. Meta has produced a lot of great models (LLama), maybe this is a comeback... but I'm cautious, as the jump in the quality is almost too high. Also, I think people aren't used that using such models requires meta.ai or meta ai app.

Given llama 4 mucked up benchmark numbers, I’d take spark announcement with a many grains of salt.

Re: Muse Spark: Scaling towards personal superintelligence

#234
Comes impressively close to GPT 5.4 / Gemini 3.1 Pro / Opus 4.6! Mostly behind OpenAI on coding/agentic benchmarks, behind Google on text reasoning, behind Anthropic on Humanity's Last Exam with tools (surprisingly the only benchmark where Anthropic leads currently).

Meta hasn’t fully caught up, but they came close and I think can solidly claim to be a frontier lab again. I’d call it a 3.5 horse race right now, and hopefully their next model improves. More model competition is good!

Poor Grok 4.2 should probably be dropped from the table.

Re: Muse Spark: Scaling towards personal superintelligence

#235

This really reinforces the idea that the AI race and the Railroad Mania of the 19th century are very similar. So many different companies are going to have similarly powerful ai that there will be no moat around it and it will be cheap. They will never earn their investment back.

Nah. Everybody is talking about ai. Everybody is using it. It's by far the most popular new tool human beings are using currently. As popular as mobile phones or spoons. And maybe as disruptive as the steam engines. AI companies are becoming the largest software companies on the planet. Everything points into that direction. Trillions of dollars are waiting in the market to be collected.

Re: Muse Spark: Scaling towards personal superintelligence

#236
I am already somewhat concerned with companies like Anthropic and especially OpenAI having personal data via chats. Typing that sort of information into a Meta AI product feels completely irresponsible. You could make some very sophisticated ads/psyop attacks with data from daily ai chats.

I doubt its better than Opus and even if it was its not worth the privacy concerns.

Re: Muse Spark: Scaling towards personal superintelligence

#238
post #2

"Muse Spark is available now, and Contemplating mode will be rolling out gradually in meta.ai." How does one get their hands on these models? They are not open-source, right? I go to meta.ai, but it's just a chat interface---no equivalent to codex or claud code? Can you use this through OpenCode? Is meta charging for model access, or is the gathering of chat data a sufficiently large tithe?

That would be my question also. I like it when companies have easy to sign up for, pay as you go models. Being able to buy $5 worth of tokens and get an API key - in less than a few minutes - is ideal.

Re: Muse Spark: Scaling towards personal superintelligence

#239
post #70
post #46

Ran some of my internal benchmarks against this and I'm very unimpressed. I don't think this moves them into the OAI v Anthropic v Gemini conversation at all. Major analytical errors in their response to multiple of my technical questions.

Playing with this some more and it's actively not good. Just basic mathematical errors riddling responses. Did some basic adversarial testing where its responses are analyzed by Gemini and Gemini is finding basic math errors across every relatively (relative to Opus, Gemini or GPT can handle) simple ask I make. Yikes.

Post actual results, make a blog post. Don't just say "this sucks" without tangible evidence.

Otherwise you're doomed to "sample size of one" level of relevance.

Post reply on HN