Live data from Hacker News

Muse Code and Muse Spark 1.2

research.meta.ai

81–90 of 268 posts

Re: Muse Code and Muse Spark 1.2

#82
post #58

Earlier quoted context omitted.

Or they spent time optimizing their model to real world problems they're facing and didn't waste time trying to game a benchmark.

Or they did try to game the benchmarks and just didn’t do it well enough. Benchmarks are one data point, not the only one, but the easiest one to compare.

Right, but the point is that you can't conclude that a model is necessarily bad because it's not hitting the same scores on benchmarks. I just don't agree with lacker's conclusion, because their logic doesn't seem to consider that. Scoring lower on a benchmark doesn't strictly mean they have a bad model, but it may be the case. Like you said it's one data point, but being the easiest, and obviously most gamed, means you should probably weigh them less heavily.

Re: Muse Code and Muse Spark 1.2

#83
post #50

They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it. They left Opus in and got beat in all but one benchmark. Nothing wrong with trying to improve, but why the marketing games? Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly. Then wh…

We can throw benchmarks in the bin by now. Each one I've seen is heavily biased and skewed. It holds very little reliable data points (unfortunately)

If you look at papers on benchmarks, they're usually created to expose gaps in how models are trained. It should be no surprise that models get better on them over time, because you can't get better at what you don't measure.

Cherry picking the benchmarks you present is where the falsehoods lie.

Re: Muse Code and Muse Spark 1.2

#85
post #70

Meta is offering a 10x discount on input ($0.10 vs. $1.25/Mtok) and 20x discount on output ($0.20 vs. $4.25/Mtok) if you opt in to let them train on your data. https://developer.meta.com/ai/models/muse-spark/

Probably to compete with DeepSeek, which AFAIK also retains data (or at least OpenRouter says they do)

FYI: there are providers of deepseek that offer the same or lower pricing and zero retention policies.

Re: Muse Code and Muse Spark 1.2

#86
post #76

Earlier quoted context omitted.

DeepSeek is really crazy cheap, though, and they don't have a giant pool of other invasive personal data to correlate it with.

I hate to say this and this is because I fucking despise meta. But between DeepSeek and Meta, and trust they handle the training data correctly, I trust meta.

What do you mean by "correctly"?

Re: Muse Code and Muse Spark 1.2

#87
post #74

Meta is offering a 10x discount on input ($0.10 vs. $1.25/Mtok) and 20x discount on output ($0.20 vs. $4.25/Mtok) if you opt in to let them train on your data. https://developer.meta.com/ai/models/muse-spark/

I actually really like that pricing strategy. It's very transparent

I love the idea of this pricing strategy but there is no way meta is not training on your data regardless of your monthly invoice

Re: Muse Code and Muse Spark 1.2

#88
post #70

Earlier quoted context omitted.

Probably to compete with DeepSeek, which AFAIK also retains data (or at least OpenRouter says they do)

FYI: there are providers of deepseek that offer the same or lower pricing and zero retention policies.

Unfortunately, none with the same caching performance as DeepSeek proper.

Re: Muse Code and Muse Spark 1.2

#89
Hey guys I'm just wondering. Usually when someone announces a new model, they'll show you some fancy viz/video/images: "These are what my model can produce." I'm wondering if anyone is keeping track of these? Like in a gallery form, "Use this prompt to produce this output".

By itself is useful ("I want something like this, I'll just reuse the prompt and tweak"), but it can also be used as a "draw me a pelican on a bicyle" alternative. Basically feeding those prompts over model releases.

Re: Muse Code and Muse Spark 1.2

#90
post #59

Here's the Muse Spark 1.2 pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... I think it's a bit of an improvement on the Spark 1.1 pelican: https://simonwillison.net/2026/Jul/9/muse-spark-1-1/

I wonder if at this point the labs ha.. please Simon, stop this bs everytime. Its time.

This one is actually a pretty good case for the continuing value of the benchmark, because it lets me visually compare Spark (8th April), Spark 1.1 (9th July), and Spark 1.2 (5th August): https://bsky.app/profile/simonwillison.net/post/3mseqv5z4qk2...
Post reply on HN