Live data from Hacker News

Muse Spark 1.3

developer.meta.com

311–320 of 475 posts

Re: Muse Spark 1.3

#312
post #5
post #3

llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle" https://tools.simonwillison.net/markdown-svg-renderer?url=ht... 4.2266 cents, 38 seconds. For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat. UPDATE: Here's an…

Is there a reason these pelicans always have roughly the same composition (side-view, 2d, biking right, flat ground beneath, etc)? I don't see any of that detailed in the prompt, yet they all seem to generate roughly the same image of differing quality.

Yes. It's because you are asking it to generate an image of a pelican riding a bicycle. If someone asked you to draw a pelican riding a bycycle, would you interpret that to mean using 3d photorealism? LLMs follow conventions. The convention for an animal riding a bike is to create a childish 2d line drawing.

Re: Muse Spark 1.3

#313
post #45

Earlier quoted context omitted.

With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.

and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments. This technology is strictly an extractive parasite on the world. Use it, bu…

Talking as if you are not disposable. If you are let go from your company, you can be easily replaceable.

People already started using contributor API, and your input is irrelevant.

Re: Muse Spark 1.3

#314
post #299

Earlier quoted context omitted.

What do you think you're doing when you accept an edit, press thumbs up, or don't ask for modifications after an edit.

Thats not exactly 'validated'. Feels very noisy, it is not a good bar for either - does this code do what the user actually asked - is this code actually 'good' There would be so many examples of coding projects that these models began or attempted to work in, that were abandoned because the models were floundering. I would imagine the labs have some decent ways to produce novel requirements and then actually validat…

This is exactly what RLVR is, and the reason that models have improved so much at verifiable domains like coding and math while not so much on unverifiable ones like writing and UI design.

Re: Muse Spark 1.3

#315
> /taste: an anti-slop filter: a flat checklist of visual defaults not to use, so generated UI stops looking machine-made.

This is interesting

Re: Muse Spark 1.3

#317

I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap and was actually really pleasantly surprised with it. It's not a frontier model by any means, but for work that didn't require a top of the line model, I really enjoyed using it. I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What…

Would be cool if there was a benchmark to evaluate the “tool-like” quality of a model - its capability to quickly, cheaply, accurately, do exactly as it is asked.

Re: Muse Spark 1.3

#318

Earlier quoted context omitted.

when are we going to stop pretending these benchmarks have any meaning? anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.

+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks. (I'm not happy about the above being true, but it's the reality I seem…

I have been using glm 5.3 flash and it feels as good as opus 5. Put a lot of work into it this week (100m tokens). Now I'm curious to try this one. These smaller models are getting very good imo

Re: Muse Spark 1.3

#319

Earlier quoted context omitted.

and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments. This technology is strictly an extractive parasite on the world. Use it, bu…

My labor makes other people's lives better, so I would expect something that replaces my labor to do the same.

You’d expect that, wouldn’t you? But, alas…
Post reply on HN