I would like someone to tell me how stupid I am. If I were Meta/Zuck I'd open source a great model the moment my company developed it. This just looks like a pitch to investors, otherwise.
"This just looks like a pitch to investors" The goal of public companies is generally to generate profit for their investors.
Muse Spark: Scaling towards personal superintelligence
251–260 of 392 posts
Re: Muse Spark: Scaling towards personal superintelligence
#252Genuine question: Why release this the day after Mythos? It does not appear SOTA (just based on benchmarks). OpenAI will likely release Spud tomorrow.
Re: Muse Spark: Scaling towards personal superintelligence
#253This really reinforces the idea that the AI race and the Railroad Mania of the 19th century are very similar. So many different companies are going to have similarly powerful ai that there will be no moat around it and it will be cheap. They will never earn their investment back.
What people seem to miss is that they don't need to get the investment back from people, they will get it from machines.
Re: Muse Spark: Scaling towards personal superintelligence
#254Do they mean "the chain of thought is visible to the user" (ie. not hidden like ChatGPT), or "the medium of the chain of thought is not text, but visuals" (ie. thinking in images).
I'd guess the former, since it wouldn't be economical to generate transient images, just for thinking. But I'm not sure why they'd highight that in that case. If it were the second thing, that'd be extremely interesting. The first model not to think in text.
Re: Muse Spark: Scaling towards personal superintelligence
#255(I'm not using it as I'm not agreeing to their ad terms).
Re: Muse Spark: Scaling towards personal superintelligence
#256Comes impressively close to GPT 5.4 / Gemini 3.1 Pro / Opus 4.6! Mostly behind OpenAI on coding/agentic benchmarks, behind Google on text reasoning, behind Anthropic on Humanity's Last Exam with tools (surprisingly the only benchmark where Anthropic leads currently). Meta hasn’t fully caught up, but they came close and I think can solidly claim to be a frontier lab again. I’d call it a 3.5 horse race right now, and h…
Re: Muse Spark: Scaling towards personal superintelligence
#257I don't get the comments trashing this. If it slightly beats or even matches Opus 4.6, it means Meta is capable of building a model competitive with the leading AI company. Sure, they spent a lot of money and will have on-going costs. But how much more work would it take to turn that into a coding agent people are willing to try (and pay for) along side their usage of a collection of agents (Claude, Codex, etc)? Also…
Re: Muse Spark: Scaling towards personal superintelligence
#258Earlier quoted context omitted.
Playing with this some more and it's actively not good. Just basic mathematical errors riddling responses. Did some basic adversarial testing where its responses are analyzed by Gemini and Gemini is finding basic math errors across every relatively (relative to Opus, Gemini or GPT can handle) simple ask I make. Yikes.
Post actual results, make a blog post. Don't just say "this sucks" without tangible evidence. Otherwise you're doomed to "sample size of one" level of relevance.
Re: Muse Spark: Scaling towards personal superintelligence
#259Comes impressively close to GPT 5.4 / Gemini 3.1 Pro / Opus 4.6! Mostly behind OpenAI on coding/agentic benchmarks, behind Google on text reasoning, behind Anthropic on Humanity's Last Exam with tools (surprisingly the only benchmark where Anthropic leads currently). Meta hasn’t fully caught up, but they came close and I think can solidly claim to be a frontier lab again. I’d call it a 3.5 horse race right now, and h…
It's looking rather low on reasoning and long-range problems with the approach described. For example, even with 16 agents and compaction, the HLE score is significantly below Anthropic's Mythos. Like you, I can see the release as a net Good Thing, but apples-to-apples for each org's latest models do have Meta holding steady in the middle pack.
Re: Muse Spark: Scaling towards personal superintelligence
#260Earlier quoted context omitted.
I'll cosign what you said, simultaneously, yr interlocutor's point is also well-founded and it depresses me it's not better known and sounds so...off...due to conventional wisdom x God King Zuck's misunderstanding his own company and resulting overreaction. They beat Gemini 2.5 Flash and Pro handily on my benchmark suite. (tl;dr: tool calling and agentic coding). Llama 4 on Groq was ~GPT 4.1 on the benchmark at ~50%…
Your comments seem to imply the engineers made a great product but Zuck intervened so now it's shit