Live data from Hacker News

Muse Code and Muse Spark 1.2

research.meta.ai

61–70 of 265 posts

Re: Muse Code and Muse Spark 1.2

#61
post #58
post #56

Earlier quoted context omitted.

My conclusion is the opposite. If benchmarks were meaningless, surely Meta would be able to find some benchmark that shows they are better than Sol and Fable. The fact that they can't do that tells me that benchmarks still do mean something.

Or they spent time optimizing their model to real world problems they're facing and didn't waste time trying to game a benchmark.

Or they did try to game the benchmarks and just didn’t do it well enough.

Benchmarks are one data point, not the only one, but the easiest one to compare.

Re: Muse Code and Muse Spark 1.2

#62
post #59

Here's the Muse Spark 1.2 pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... I think it's a bit of an improvement on the Spark 1.1 pelican: https://simonwillison.net/2026/Jul/9/muse-spark-1-1/

Lol is that a helmet? Clearly they're taking the Anthropic path and are safety pilling their models

Re: Muse Code and Muse Spark 1.2

#63

Will someone at Meta for the love of God make it so none of this stuff goes through Facebook.com? You want customers but most corporate firewalls block social media. Also, a lot of devs do not want their work stuff tied up to their facebook account. For the love of all things show the IG / FB logins as optional and do email as primary. I am not a fan of Meta but I do cheer for any competitors against OpenAI and Anthr…

I could sign in with my Meta account that is independent and not linked to IG or Facebook. Just click "Login with Email" on dev.meta.ai.

Re: Muse Code and Muse Spark 1.2

#64

They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it. They left Opus in and got beat in all but one benchmark. Nothing wrong with trying to improve, but why the marketing games? Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly. Then wh…

> They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it.

If you scroll very slightly farther there is a benchmark that includes Sol, showing it outperforming Terra (as expected) and Spark 1.2

Re: Muse Code and Muse Spark 1.2

#65
post #4

Earlier quoted context omitted.

Clearly they're positioning it as a mid model.

Which is in itself a bit weird as mid models nowadays are a golden mean fallacy. Terra is much less popular than both Luna (cost-sensitive) and Sol (performance-sensitive). Claude Sonnet is a weird exception to the mid models because Anthropic doesn't do much with Haiku and Opus is too big.

I don’t know where you get your statistics, but I love Terra and use it all of the time. It is the default fastest model in ChatGPT/Codex today. I saw today a notice saying that the model had hit capacity briefly

Re: Muse Code and Muse Spark 1.2

#66

Muse Spark 1.1 was released July 16th, less than a month ago. A new version release this soon (particularly after Kimi K3's release drastically overshadowed it) is a bit sus and it appears that Meta is trying a first launch do-over.

Doesnt seem suspect to me, training runs have checkpoints and there is no reason you cant release a checkpoint even if you are still training the model

Re: Muse Code and Muse Spark 1.2

#69
post #60
post #26

Last I heard, everyone at Meta was using Claude Code. Any insiders know how Muse Code is doing internally?

Everyone is still using claude or codex if they aren’t forced off of it. Nobody is going to use a worse tool in this culture.

it's actually interesting that they're not being forbidden to use claude/codex, is Meta paying for it or is it personal accounts?

Re: Muse Code and Muse Spark 1.2

#70

Meta is offering a 10x discount on input ($0.10 vs. $1.25/Mtok) and 20x discount on output ($0.20 vs. $4.25/Mtok) if you opt in to let them train on your data. https://developer.meta.com/ai/models/muse-spark/

Probably to compete with DeepSeek, which AFAIK also retains data (or at least OpenRouter says they do)
Post reply on HN