Live data from Hacker News

Muse Spark 1.3

developer.meta.com

291–300 of 475 posts

Re: Muse Spark 1.3

#291
post #5

Earlier quoted context omitted.

Is there a reason these pelicans always have roughly the same composition (side-view, 2d, biking right, flat ground beneath, etc)? I don't see any of that detailed in the prompt, yet they all seem to generate roughly the same image of differing quality.

Not when rendered via POV-Ray: https://blog.nawaz.org/posts/2025/Oct/pelican-on-a-bike-rayt... I plan to update it with more pelicans from all the models released since. (Spoiler alert: They haven't improved much since then).

Ohh, horizontal wheels. They’re about as good as I expected, models have pretty bad spatial awareness. I would expect Fable to be a bit better than old models, though.

Re: Muse Spark 1.3

#292
post #5

Earlier quoted context omitted.

Is there a reason these pelicans always have roughly the same composition (side-view, 2d, biking right, flat ground beneath, etc)? I don't see any of that detailed in the prompt, yet they all seem to generate roughly the same image of differing quality.

Not when rendered via POV-Ray: https://blog.nawaz.org/posts/2025/Oct/pelican-on-a-bike-rayt... I plan to update it with more pelicans from all the models released since. (Spoiler alert: They haven't improved much since then).

I wonder how a multi-modal model would do with a harness and tool calling? Specifically a "render" command that produced an image output enabling it to iterate. (Well I see you did this manually with gemini 2.5 pro but I still think it would be interesting to explore various harness setups.)

> GPT-5.1 Codex

> monstrosity

What are you talking about? That's clearly a sci-fi pelican on a hoverboard (successor of the humble bicycle) wearing a visor. Truly visionary.

Re: Muse Spark 1.3

#293
post #78

Earlier quoted context omitted.

I sincerely believe I've never had a single original thought™ in my whole life. There is this scene in the HBO series Westworld where a "host" says some words in sequence which is shown on a display as she says it. Of course, even me thinking of this scene and connecting it to your comment was not original, someone else clearly had the same programming as me. A medium blog post says > Pair what with me?” — the moment…

Westworld is such a time capsule. It's not even that old - but back when it was aired, an AI that can not just string together coherent sentences, but produce coherent reactions in novel, fully unintended contexts, like Maeve was doing there? It was totally a sci-fi premise. Now we have AIs capable of that and more, and no one bats an eye.

Indeed: “Our hosts began to pass the Turing test within the first year.”

Required sci-fi suspension-of-disbelief in 2017, and then at some point in the last few years we just blew by that one.

Later seasons of the show were much less dramatically satisfying, but also played out the consequences of the science of artificial intelligence demonstrating as a side-effect that human intelligence and free will might have as much of an uncertain foundation as that of machines.

How much data from the Panopticon, how many parameters would it take to train a model that could predict your responses?

Re: Muse Spark 1.3

#295
post #230

Earlier quoted context omitted.

and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments. This technology is strictly an extractive parasite on the world. Use it, bu…

I'm using AI to build things I wouldn't (and/or couldn't) have built before. That's the opposite of parasitic.

[flagged]

Re: Muse Spark 1.3

#296
post #213
post #122

Earlier quoted context omitted.

It would be very embarrassing for any lab to benchmaxx the pelican on bicycle svg prompt, since it would be very easy to detect it by varying the prompt.

The amount of discussion around it means that the test and all the reviews of results, images, approaches etc are implicitly included in training data. It’s not deliberate “benchmaxxing” but things that are discussed a lot online are naturally things that LLMs learn better.

You can't benchmaxx spatial awareness without solving the fully general problem (at least I figure).

Re: Muse Spark 1.3

#297
post #236

Earlier quoted context omitted.

Every lab trains their models with AI generated code at this point.

Hopefully, 'validated' AI code

What do you think you're doing when you accept an edit, press thumbs up, or don't ask for modifications after an edit.

Re: Muse Spark 1.3

#298
post #45

Earlier quoted context omitted.

With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.

and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments. This technology is strictly an extractive parasite on the world. Use it, bu…

Don’t you have some looms to break?

Re: Muse Spark 1.3

#299
post #236

Earlier quoted context omitted.

Hopefully, 'validated' AI code

What do you think you're doing when you accept an edit, press thumbs up, or don't ask for modifications after an edit.

Thats not exactly 'validated'. Feels very noisy, it is not a good bar for either - does this code do what the user actually asked - is this code actually 'good'

There would be so many examples of coding projects that these models began or attempted to work in, that were abandoned because the models were floundering.

I would imagine the labs have some decent ways to produce novel requirements and then actually validate they are met, without the noisiness of implicit human feedback.

That said, the more I think about it, you are right, there's probably also very good ways to extract signal for all these sessions.

Re: Muse Spark 1.3

#300
post #273
post #250

Earlier quoted context omitted.

Any more info on this?

Cline experiment: https://x.com/cline/status/2085237843379519737 Muse code: https://developer.meta.com/ai/resources/blog/build-with-muse... > Co-trained with the harness. Muse Code was in the training loop from day one, so tool calls succeed and plans execute cleanly. Crucially, we trained across multiple harnesses, so while the model is at its best in Muse Code, it still generalizes to other coding agents you alread…

thank you
Post reply on HN