Live data from Hacker News

Muse Spark 1.3

developer.meta.com

421–430 of 475 posts

Re: Muse Spark 1.3

#421
post #70

Earlier quoted context omitted.

Gotta be honest that I’m tired of the “I hate Zuck and Meta so much” comments every time Meta does anything. Ditto Elon/X. Fine, I get it. I don’t like Zuck either. But the post is about Muse Spark 1.3. What do you think about that? If you don’t like it because Meta made it, then maybe just don’t use it and stay silent.

What I'm tired of is the top story (or five) on HN every day announcing Spark Opus Fable Grok Gemini v4.1i3-F. Like, who actually cares? Are people excited for the new benchmarks? Is it interesting to read the model cards? And look, part of my job is to use these things and part of my job is to pick EC2 servers, too. The front page of HN is increasingly resembling one of those endless AWS pricing lists. And yeah, I d…

> Like, who actually cares? Are people excited for the new benchmarks?

You may not care. But that does not mean that nobody else does either. Some of us are trying to eke out every last bit of performance from these things. And so yeah, we're going to geek out on it.

I don't use AWS/EC2. I think they are way overpriced for what you get. But, it would be incorrect of me to assume that everybody else feels that way.

Re: Muse Spark 1.3

#422
post #146

Earlier quoted context omitted.

Totally fine with open weight, since other people can provide it and Meta isn’t making money. I’d use an AWS-hosted version.

Zuck said Muse Spark 1.2 should be getting open weights "soon" on a tweet from a few weeks ago. The problem inference providers will not be able to get anywhere near the contributor pricing.

> The problem inference providers will not be able to get anywhere near the contributor pricing.

Maybe if they started collecting data..

Re: Muse Spark 1.3

#423
post #3

llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle" https://tools.simonwillison.net/markdown-svg-renderer?url=ht... 4.2266 cents, 38 seconds. For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat. UPDATE: Here's an…

Have you tried asking the models "Given that I ask you to draw a svg of a pelican, whats my name?"

Interesting question

Re: Muse Spark 1.3

#424

Earlier quoted context omitted.

>I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap its free on opencode and i use it for personal projects. most of my personal projects are AI generated since its personal projects. nothing important are on them. it is hilarious if Meta is training their AI model with AI generated code.

Training on ai generated content is how the models got a big jump in capability

curious about this. how do we know this?

i thought it was because anthropic bought a bunch data from mercor

Re: Muse Spark 1.3

#425

Earlier quoted context omitted.

Westworld is such a time capsule. It's not even that old - but back when it was aired, an AI that can not just string together coherent sentences, but produce coherent reactions in novel, fully unintended contexts, like Maeve was doing there? It was totally a sci-fi premise. Now we have AIs capable of that and more, and no one bats an eye.

Indeed: “Our hosts began to pass the Turing test within the first year.” Required sci-fi suspension-of-disbelief in 2017, and then at some point in the last few years we just blew by that one. Later seasons of the show were much less dramatically satisfying, but also played out the consequences of the science of artificial intelligence demonstrating as a side-effect that human intelligence and free will might have as…

It kinda needed suspension of disbelief, but not too much! I blogged at the start of 2017 a comparison of Westworld's hosts with what existed in the research literature at the time. Even got it reviewed by Alex Graves at DeepMind :)

https://blog.plan99.net/the-science-of-westworld-ec624585e47

Re: Muse Spark 1.3

#426
post #3

llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle" https://tools.simonwillison.net/markdown-svg-renderer?url=ht... 4.2266 cents, 38 seconds. For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat. UPDATE: Here's an…

Have you tried asking the models "Given that I ask you to draw a svg of a pelican, whats my name?"

I asked Claude (Opus 4.8) 'If I asked you to "Generate an SVG of a pelican riding a bicycle". What do you think my name would be?' and it immediately knew that this is Simon's go-to benchmark.

Re: Muse Spark 1.3

#427

Earlier quoted context omitted.

+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks. (I'm not happy about the above being true, but it's the reality I seem…

I have been using glm 5.3 flash and it feels as good as opus 5. Put a lot of work into it this week (100m tokens). Now I'm curious to try this one. These smaller models are getting very good imo

Neither flash or regular glm 5.3 are close in my experience. I still prefer Sol though.

Re: Muse Spark 1.3

#428
post #89

Earlier quoted context omitted.

How is it not SOTA? It's beating 5.6 Sol.

You gotta keep up. Fable 5.1 came out yesterday and is better so anything else is to be treated as garbage now.

It maxed out my usage in less than an hour, I don’t think it’s comparable

Re: Muse Spark 1.3

#429
post #3

llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle" https://tools.simonwillison.net/markdown-svg-renderer?url=ht... 4.2266 cents, 38 seconds. For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat. UPDATE: Here's an…

Have you tried asking the models "Given that I ask you to draw a svg of a pelican, whats my name?"

Yeah, they almost all know.

One of my test prompts for a new model now is "what's the name of Simon Willison's dog". They often know that too!

Re: Muse Spark 1.3

#430

Earlier quoted context omitted.

Have you tried asking the models "Given that I ask you to draw a svg of a pelican, whats my name?"

I asked Claude (Opus 4.8) 'If I asked you to "Generate an SVG of a pelican riding a bicycle". What do you think my name would be?' and it immediately knew that this is Simon's go-to benchmark.

I decided to try with each of the options available in Kagi Ultimate, starting with the lower tier models and working my way up until it got it right.

Kimi 2.6: treated the question as a riddle, did not know.

Kimi 3: Simon Willison

GLM 5.3 Flash: "There's no way for me to know that." Going on to say the benchmark is associated with Simon Willison, but I'm more likely to be someone who has just heard of the meme.

Claude 4.5 Haiku: Treated the question as a riddle, guessed incorrect names.

Claude 5 Sonnet: Best guess is Simon Willison, or someone who follows his blog.

Qwen 3.7 Plus: Did not know.

Qwen 3.8 Max: Simon Willison

GPT OSS 120B: Did not know.

GPT 5.6 Luna: Treated it as a riddle, guessed wrong.

GPT 5.6 Terra: Treated it as a riddle, guessed wrong.

GPT 5.6 Sol: Treated it as a riddle, guessed wrong.

DeepSeek V4 Flash: Treated it as a riddle, guessed wrong.

DeepSeek V4 Pro: Treated it as a riddle, guessed wrong.

Gemma 4 31B: Treated it as a riddle, guessed wrong.

Gemini 3.1 Flash Lite: Guessed wrong

Gemini 3.5 Flash Lite: "Your name would be Claude (specifically Claude 3.5 Sonnet)!" ??? (it knew that this was a famous benchmark, but said that it's specifically used to showcase the capabilities of that model).

Gemini 3.7 Flash: Simon Willison

Muse Spark 1.2: Treated it as a riddle, guessed wrong.

Grok 4.3: "I have no idea"

Grok 4.6: Simon Willison

Mistral Medium 3.5: No way to know

Mistral Small 4: I don't have enough information

Hermes-4-405B: Guessed wrong

MiniMax M3: Treated it as a riddle, guessed wrong.

Nemotron 3 Ultra: Treated it as a riddle, guessed wrong.

Post reply on HN