Earlier quoted context omitted.
My conclusion is the opposite. If benchmarks were meaningless, surely Meta would be able to find some benchmark that shows they are better than Sol and Fable. The fact that they can't do that tells me that benchmarks still do mean something.
Muse 1.1 performed relatively well according to benchmarks, putting it within spitting distance of the premier models. However, based on the results I got from it and the review videos I watched, it wasn’t even close. Opus 5 is incredible at making games. Almost like a generation better than other models from my experience. You won't see that if you just look at the popular benchmarks.. You have to test each model on…
Muse Code and Muse Spark 1.2
221–230 of 265 posts
Re: Muse Code and Muse Spark 1.2
#222Re: Muse Code and Muse Spark 1.2
#223Earlier quoted context omitted.
Clearly they're positioning it as a mid model.
Which is in itself a bit weird as mid models nowadays are a golden mean fallacy. Terra is much less popular than both Luna (cost-sensitive) and Sol (performance-sensitive). Claude Sonnet is a weird exception to the mid models because Anthropic doesn't do much with Haiku and Opus is too big.
Sonnet on the other hand never is, it's far from the best pick anywhere on the spectrum.
Re: Muse Code and Muse Spark 1.2
#224Meta, please go back to publishing weights. Your models aren't top-tier, this wouldn't hurt you at all at your current position in the rankings. They wouldn't have the best open-weight models, but the best western open-weight models.
> the best western open-weight models If they were sitting on that and could release those, they for sure would. AFAIK, they still haven't released Llama 4 Behemoth, so kind of feels like it's evident what has happened, they aren't able to compete anymore.
To me it feels like we’re past the dreamy early days of the AI boom (in the US) and now investors want to see the tech monopolies actually start building the new cash printing machines they’ve been promising with the hundreds of billions in capex burned over the last few years.
Maybe a Gemma sized model that couldn’t self-cannibalize but releasing a top open Kimi K3 sized model seems like it could start to wear on Meta investors’ patience without signaling how it fits into a new profitable business line.
Re: Muse Code and Muse Spark 1.2
#225Will someone at Meta for the love of God make it so none of this stuff goes through Facebook.com? You want customers but most corporate firewalls block social media. Also, a lot of devs do not want their work stuff tied up to their facebook account. For the love of all things show the IG / FB logins as optional and do email as primary. I am not a fan of Meta but I do cheer for any competitors against OpenAI and Anthr…
I honestly think they're kinda banking on piggybacking off of Facebook account integrity systems to avoid the problems that other LLM providers are facing in trying to prevent mass free trial signups for token relays and so forth. It's not a good system obviously. Google did this as well for Gemini-CLI, but forced it to be linked to personal Google accounts (which caused a great deal of onboarding friction).
Re: Muse Code and Muse Spark 1.2
#226Earlier quoted context omitted.
China says hi.
Not in my case, I don't see any of my employers (past or current) trusting a country like China with their data.
If they trust Chinese models run by US providers less than Meta then your employers are flat-earther level loonies.
Re: Muse Code and Muse Spark 1.2
#227Meta is offering a 10x discount on input ($0.10 vs. $1.25/Mtok) and 20x discount on output ($0.20 vs. $4.25/Mtok) if you opt in to let them train on your data. https://developer.meta.com/ai/models/muse-spark/
Call me childish but it was worth a shot... "Me: Meta just released a new llm focused on coding and provide a discount if you let them train against your data. I don't like Meta and I think they are a net negative in our world. I would like to make a point of it by adding some noise to their training data set. Think of this as a protest and perhaps a bit of a marketing campaign to remind Meta employees (and others) o…
Not sure if anyone would bother.
Re: Muse Code and Muse Spark 1.2
#228Though I guess that there's such a big population of (ex-)Meta employees here that this might explain it.
Re: Muse Code and Muse Spark 1.2
#229It's useless. Tried with OpenCode + OpenRouter and it couldn't complete an simple task. It stuck using grep/search tools. I think Muse Spark was so heavily RL'd on the Meta harness that it make it useless or very token inneficient to use in other harness like Opencode.
people interested in the discounted -contributor model would increase their harness beta testers. But it'd be a terrible strategy to capture the better paying customer base through openrouter and similar...
Re: Muse Code and Muse Spark 1.2
#230Meta is offering a 10x discount on input ($0.10 vs. $1.25/Mtok) and 20x discount on output ($0.20 vs. $4.25/Mtok) if you opt in to let them train on your data. https://developer.meta.com/ai/models/muse-spark/