Live data from Hacker News

Muse Spark 1.3

developer.meta.com

281–290 of 475 posts

Re: Muse Spark 1.3

#281
post #32

DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap! Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!

Gemini 3.8 flash has better rates. $0.75 per million input tokens and $3.75 per million output tokens.

Compare that to Muse spark 1.3

$1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing)

It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.

Re: Muse Spark 1.3

#282

Earlier quoted context omitted.

and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments. This technology is strictly an extractive parasite on the world. Use it, bu…

I’m retired so it won’t be replacing my labor :)

The sibling reply to this is just such lazy thinking, such a trite cliche. Yes, all members of a generation are bad, end of story. Can we get back to the war between the sexes now?

Re: Muse Spark 1.3

#283
post #45

Earlier quoted context omitted.

With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.

and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments. This technology is strictly an extractive parasite on the world. Use it, bu…

[flagged]

Re: Muse Spark 1.3

#284

Earlier quoted context omitted.

You gotta keep up. Fable 5.1 came out yesterday and is better so anything else is to be treated as garbage now.

Look at all the valuable software products that Fable 5.1 has produced since yesterday!

If it talks less like a robot, I’d call that a win!

Re: Muse Spark 1.3

#285
post #32

DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap! Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!

when are we going to stop pretending these benchmarks have any meaning? anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.

+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks.

(I'm not happy about the above being true, but it's the reality I seem to inhabit.)

Re: Muse Spark 1.3

#286
post #3

llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle" https://tools.simonwillison.net/markdown-svg-renderer?url=ht... 4.2266 cents, 38 seconds. For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat. UPDATE: Here's an…

I see no point having these pelicans used for anything related model qualification.

Re: Muse Spark 1.3

#287

Earlier quoted context omitted.

and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments. This technology is strictly an extractive parasite on the world. Use it, bu…

My labor makes other people's lives better, so I would expect something that replaces my labor to do the same.

global development and relief of poverty has relied on there being an economic surplus for all from organized labor. everyone gets a benefit although it is unfairly distributed.

i think that there is growing organized labor today that produces no surplus. instead, it transfers wealth from some to others, causing net harm to all in the process. an example of this would be purdue pharma.

depending on who you ask the list of jobs and industries which have zero surplus is getting large. swathes of private equity and leveraged financial instruments, shitcoins, management consultancy, are pure deadweight loss.

the work does nothing or causes net harm.

Re: Muse Spark 1.3

#289
post #3

llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle" https://tools.simonwillison.net/markdown-svg-renderer?url=ht... 4.2266 cents, 38 seconds. For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat. UPDATE: Here's an…

[flagged]

"I was interviewed for a job as a software developer last week and they asked me to draw a picture of a pelican riding a bicycle. Aced it, got the job as a senior software engineer."

That is the best joke I have heard this year. Ready for a stand-up comedy special. Or a song. Superb!

Re: Muse Spark 1.3

#290

Earlier quoted context omitted.

I'm wondering whether anyone has yet extracted AWS keys from a model trained on user input. Because users are definitely feeding secrets into these "contributor" models

doesn't mean the raw text goes into training. they most likely have a pipeline to clean out any secrets before they train on it?

In theory, but in practice how difficult is that?
Post reply on HN