Live data from Hacker News

Muse Code and Muse Spark 1.2

research.meta.ai

51–60 of 265 posts

Re: Muse Code and Muse Spark 1.2

#52

Meta is offering a 10x discount on input ($0.10 vs. $1.25/Mtok) and 20x discount on output ($0.20 vs. $4.25/Mtok) if you opt in to let them train on your data. https://developer.meta.com/ai/models/muse-spark/

This makes it a very interesting alternative to Deepseek for personal work where I don't care about the training - judging by the AA benchmarks it seems like overall cost per task is similar to the new Deepseek Flash but with better benchmarks (and inbuilt vision capabilities).

Re: Muse Code and Muse Spark 1.2

#53

They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it. They left Opus in and got beat in all but one benchmark. Nothing wrong with trying to improve, but why the marketing games? Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly. Then wh…

While I won't take their limited benchmarks with much salt, if it actually is this close to opus, but at a third the cost, that's pretty solid. Now, Terra is pretty damn affordable too and you're right that it's suspicious that they don't put Sol in there at all.

Re: Muse Code and Muse Spark 1.2

#54
post #48
post #40

Somewhat surprised that Meta with all their resources couldn’t make a model that matches Composer on any frontier. All the Sparks are dominated by some other model everywhere along the frontier. Nothing fancy here since Llama defined the open model. The use traces must be crucial to functionality which is why they’re keeping prices so low.

They rebooted less than one year ago so this is decent progress. Obviously users don't care about progress though.

Yeah, progress is useful as an internal metric, but I'm going to measure against the present frontier unfortunately. Eager to see what they come up with in the future.

Re: Muse Code and Muse Spark 1.2

#55

Earlier quoted context omitted.

If there were, do you believe it would be in their interest to answer this publicly?

> > Any insiders know how Muse Code is doing internally? > If there were, do you believe it would be in their interest to answer this publicly? If it were being adopted like gangbusters in their organization, sure! So... the fact that nobody is volunteering the information is probably a valid signal of how things are actually going...

You should play Blood on the clocktower

Re: Muse Code and Muse Spark 1.2

#56
post #50

They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it. They left Opus in and got beat in all but one benchmark. Nothing wrong with trying to improve, but why the marketing games? Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly. Then wh…

We can throw benchmarks in the bin by now. Each one I've seen is heavily biased and skewed. It holds very little reliable data points (unfortunately)

My conclusion is the opposite. If benchmarks were meaningless, surely Meta would be able to find some benchmark that shows they are better than Sol and Fable. The fact that they can't do that tells me that benchmarks still do mean something.

Re: Muse Code and Muse Spark 1.2

#57
post #13

This is a nice release and a solid improvement over Spark 1.1. It compares favorably with Grok 4.5. Not SOTA, but solid releases. I think they need to really get this more competitive with Deepseek V4 Flash / Luna pricing to move the needle.

If you are happy to share data for training, the contributor mode offers amazing price $0.10 / $0.20

Yes, that is the really compelling thing here IMO. Its a viable deepseek competitor for many people, and I missed that on the first pass.

Re: Muse Code and Muse Spark 1.2

#58
post #56
post #50

Earlier quoted context omitted.

We can throw benchmarks in the bin by now. Each one I've seen is heavily biased and skewed. It holds very little reliable data points (unfortunately)

My conclusion is the opposite. If benchmarks were meaningless, surely Meta would be able to find some benchmark that shows they are better than Sol and Fable. The fact that they can't do that tells me that benchmarks still do mean something.

Or they spent time optimizing their model to real world problems they're facing and didn't waste time trying to game a benchmark.

Re: Muse Code and Muse Spark 1.2

#60
post #26

Last I heard, everyone at Meta was using Claude Code. Any insiders know how Muse Code is doing internally?

Everyone is still using claude or codex if they aren’t forced off of it. Nobody is going to use a worse tool in this culture.
Post reply on HN