Live data from Hacker News

Muse Spark 1.3

developer.meta.com

301–310 of 475 posts

Re: Muse Spark 1.3

#301

Earlier quoted context omitted.

Progress is iterative. Everyone is always riffing on other’s ideas and can execute on them given enough support (eg $$). The person to get to an idea first is just 5% away, so it’s possible to catch up. Moreover,I think it’s impossible to know if you’re hitting a portion of the sigmoid, because there will often be an idea that changes the trajectory altogether. In 2024, there was a ton of talk about the plateau. Reas…

o1 was first, and Anthropic were doing a bit of it; DeepSeek brought it to the masses, but did not invent it.

Totally, RLVR as a concept predates DeepSeek; but they proposed a version that was simple and scalable. Popularizing a specific version of a technique is exactly what I mean by iterations on a theme. It’s only 5% different from what others tried before, but that 5% difference showed a lot more potential than other versions of the same idea.

Since DeepSeeks GRPO, they’ve been improvements as well like AliBabas GSPO that have gotten wide adoption. Again iterations

Re: Muse Spark 1.3

#302
For folks who are impressed with costs, why does it matter to you? Is subscriptions not a thing? I may be missing something but only companies should really care about this I would think?

Re: Muse Spark 1.3

#303

Earlier quoted context omitted.

when are we going to stop pretending these benchmarks have any meaning? anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.

+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks. (I'm not happy about the above being true, but it's the reality I seem…

And Fable 5.x makes Opus 5 look pretty dim, despite benchmarks suggesting they're comparable. The benchmarks really are just kinda meaningless.

Re: Muse Spark 1.3

#304
post #37

Earlier quoted context omitted.

Progress is iterative. Everyone is always riffing on other’s ideas and can execute on them given enough support (eg $$). The person to get to an idea first is just 5% away, so it’s possible to catch up. Moreover,I think it’s impossible to know if you’re hitting a portion of the sigmoid, because there will often be an idea that changes the trajectory altogether. In 2024, there was a ton of talk about the plateau. Reas…

> Deepseek proposes RLVR as a way to get around the lack of $ they have to produce human reasoning trace data. What was the difference between what deepseek did for R1 and what OpenAI did for o1?

openai did human crafted chain of thought dataset training. deepseek didn't have the resources so they attempted RL. doing RL correctly is hard because of the risk of model collapsing.

Re: Muse Spark 1.3

#305

For folks who are impressed with costs, why does it matter to you? Is subscriptions not a thing? I may be missing something but only companies should really care about this I would think?

Some of us own and run companies? Cost per performance is a huge deal.

Re: Muse Spark 1.3

#306

Earlier quoted context omitted.

They don't all suck equally. Here's the order, from best to worst. Amodei Google SamA Zuck Elon

In my personal opinion (this will be controversial and feel free to disagree): Elon is the best. * great contributions to many industries including spaceflight, electric cars, and self driving cars. It doesn't even matter if he is the technical mind behind these achievements or if he is just a buffoon that pretends to know the implementation details; the dude has a way of bringing together experts, having the overall…

I'm upvoting you purely because I'm sick of comments being made in good faith getting an automatic downvote

Re: Muse Spark 1.3

#307

muse-spark-1.3-contributor. Say what you want and Meta, changing the pricing to explicitly say 'we train on this and value it this much' is what every model provider should do. As a side note, it is now completely obvious how much stealing my tokens for training is worth to model providers. I avoid/pay extra/try my best to make sure I am not getting trained on but it seems like it keeps popping up that I missed a set…

This has been my hunch for a while about all the discourse of "OpenAI/Anthropic subscription pricing is unsustainable!!"

We understand theoretically they're taking our data, but yeah, that data is vital to the entire business plan of all these companies and WAY more valuable than people are giving credit for.

I checked up on Mistral recently and saw their Claude-alike coding harness is using GLM now, whatever it takes to keep users on their platform and feeding them data.

Re: Muse Spark 1.3

#308

Earlier quoted context omitted.

[flagged]

If you could actually hand write SVG code on the spot that looked like a realistic pelican riding a bike I would want to hire you for SOMETHING.

Or placed in an asylum next to the people who designed XML

Re: Muse Spark 1.3

#309
post #121

I had no idea Meta has a coding agent harness. Does anyone have experience with it and can comment? The 1.3 contributor prices look very attractive. I'll probably start using their API if performance is good and the API is reliable with decent rate limits.

You should use their harness. They trained it on multiple harnesses but have specifically optimized it for their harness. Cline also did an independent experiment w spark 1.2 where using the native harness makes it use fewer tokens / turns to accomplish tasks

Thanks. Just downloaded and pretty impressed so far. It's fast and nice to work with.

Re: Muse Spark 1.3

#310
post #24

Practically free for "contributors" at 0.2 usd/mtok. That's going to be hard to say no to for hobbyists.

I'm wondering whether anyone has yet extracted AWS keys from a model trained on user input. Because users are definitely feeding secrets into these "contributor" models

If my experience with image generation is any indication, unless AWS keys are somehow extremely prevalent in the training data, you may get something that looks like one, but it definitely won't be valid.
Post reply on HN