Live data from Hacker News

Muse Spark 1.3

developer.meta.com

361–370 of 475 posts

Re: Muse Spark 1.3

#361

I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap and was actually really pleasantly surprised with it. It's not a frontier model by any means, but for work that didn't require a top of the line model, I really enjoyed using it. I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What…

Would be cool if there was a benchmark to evaluate the “tool-like” quality of a model - its capability to quickly, cheaply, accurately, do exactly as it is asked.

This is interesting because I have transitioned to where I use SOT models.. but I kind of use them like employees that I can delegate to. I still review code.

However, I now literally say.. "Here is my objective and here is a starting point for documentation. Research this and build up a plan."

This can be very company specific, like migration from one framework to another in house infrastructure framework. I'm spending my time figuring out how the plan should be chopped so I can have confidence in the parts and not overwhelmed. I don't want a tool, I want a model that can stitch resources together into a plan. That type of model is in a whole other ballpark.

Re: Muse Spark 1.3

#362

Earlier quoted context omitted.

In my personal opinion (this will be controversial and feel free to disagree): Elon is the best. * great contributions to many industries including spaceflight, electric cars, and self driving cars. It doesn't even matter if he is the technical mind behind these achievements or if he is just a buffoon that pretends to know the implementation details; the dude has a way of bringing together experts, having the overall…

I'm upvoting you purely because I'm sick of comments being made in good faith getting an automatic downvote

Sorry for the wrongthink. Obviously I deserve my downvote, so I can never reach the 500 karma necessary to downvote others. I don't have the right opinions. (Site guidelines: don't comment on downvotes. Yeah, I know. But I'm so sick of this culture)

Re: Muse Spark 1.3

#363

Im a caveman writing c/cpp. Last time ms1.2 was even worth than DeepSeek v4f preview on internal benchmark. It just feels like extremely over fitting on certain paths.

I have opposite result: MS1.2 wrote C code without following original source code writing style, and no descriptive info why writing such code, DS4F or even Mimo seems better to me.

I think the person you are answering to was saying the same thing. They wrote "worth" instead of "worse".

Re: Muse Spark 1.3

#364
post #3

llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle" https://tools.simonwillison.net/markdown-svg-renderer?url=ht... 4.2266 cents, 38 seconds. For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat. UPDATE: Here's an…

"The LLM is better because the pelican hat is better"

Benchmarking like never before

Re: Muse Spark 1.3

#366
post #134

I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap and was actually really pleasantly surprised with it. It's not a frontier model by any means, but for work that didn't require a top of the line model, I really enjoyed using it. I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What…

If it's a mistake, it should course-correct. I agree that some of the smarter models are actually worse. I hope they take a model that's good enough--there are many--and just try to get it chatjimmy.ai speed. I have to think that's the future, somehow, and I'm really excited about it.

"If it's a mistake, it should course-correct"

Maybe, or maybe not. The thing is, that "mistake" isn't something that is generally valid. For example, the enshittification of Google Search through the last 15 years seems to be a mistake --- but perhaps not from the money-making point of view of Google Shareholders. Likewise the enshittification of reddit --- we nerdy users see it as a mistake. But for them this intended enshittification probably increased revenue.

It's the money, always the money! PR-speak like "customer satisfaction is our highest goal" is, like most PR-speak, a blatant lie.

And so it can very well be the case that for coders the frontier models get worse, but they get better for other applications --- and that all of this is just driven by "how can we capitalize the most out of it", not satisfaction levels of programmers.

Re: Muse Spark 1.3

#367
post #5
post #3

llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle" https://tools.simonwillison.net/markdown-svg-renderer?url=ht... 4.2266 cents, 38 seconds. For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat. UPDATE: Here's an…

Is there a reason these pelicans always have roughly the same composition (side-view, 2d, biking right, flat ground beneath, etc)? I don't see any of that detailed in the prompt, yet they all seem to generate roughly the same image of differing quality.

If you look at bike product photography it's always drive side facing the camera, which means front wheel on the right. If I had to guess this is probably where this comes from

Don't know if that's ever possible to know though unless you train a model from scratch but remove all bike product photography and adjacent materials from the training data?

Re: Muse Spark 1.3

#368
post #32

DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap! Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!

when are we going to stop pretending these benchmarks have any meaning? anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.

Yep these software benches are only good at testing how well they can one shot. For the kind of attended/assisted development most of us do with agents it’s hard to find a benchmark that reflects my own experience of the frontier models still being quite far ahead.

Re: Muse Spark 1.3

#369

Earlier quoted context omitted.

[flagged]

Buddy, admitting your thought processes forcibly terminate on pre-programmed keywords isn't a flex.

Why "terminate", Marxist philosophy is a legitimate topic, deserving to be studied. Like a rich sci-fi lore or a history of Tarot magic. Deep, fascinating, and wrong.

Re: Muse Spark 1.3

#370

Earlier quoted context omitted.

[flagged]

Would you care to discuss the topic, or just throw grenades? Surely you can come up with something more substantive than this

Ok. This requires the notion of "intrinsic value", which I believe does not exist (all value is subjective), yet is a foundation of all Marxist theory.
Post reply on HN