Muse Spark 1.3
441–450 of 475 posts
Re: Muse Spark 1.3
#442muse-spark-1.3-contributor. Say what you want and Meta, changing the pricing to explicitly say 'we train on this and value it this much' is what every model provider should do. As a side note, it is now completely obvious how much stealing my tokens for training is worth to model providers. I avoid/pay extra/try my best to make sure I am not getting trained on but it seems like it keeps popping up that I missed a set…
Re: Muse Spark 1.3
#443Earlier quoted context omitted.
It's really interesting, isn't it? They almost always cycle from left to right - but I have had a few which cycle in the other direction. The 2D / flat ground feels reasonable for a SVG, which implies a vector illustration.
It's my impression that it's common in western culture, where text is read left to right, and timelines are visualized as going from left to right, to also animate things going from left to right, since westerners thus have an instinct that "right = forward", so it "feels right" (familiar). I wonder to which degree this is reflected in the training data? And if you'd be more likely to get left-facing pelicans if you…
Re: Muse Spark 1.3
#444Earlier quoted context omitted.
I think where you and I disagree is on whether Anthropic is especially trustworthy on the "safety" front, more trustworthy than various other labs, especially those that produce open models, for example. I simply don't trust Amodei more than I trust, say, Liang Wenfeng. I'm not saying I trust any of them, particularly, I am saying that if a few billionaires have access to this technology, I want access to this techno…
I think you're mischaracterizing Anthropic uncharitably and lumping them in with other, less savory tech billionaires, and also not thinking through the nuances here. First, Amodei has taken an unusually strong stance among tech companies for not supplying fascist regimes with fascist tooling; in fact, even when threatened with being labeled a national security risk unless he bent the knee, he didn't. Compare and con…
But, I'll come back to "two things can be true". Anthropic is better than some, and in some regards they are navigating a complicated ethical landscape with more care than others. On the other hand, it really looks like they're angling to regulate their open competitors out of the game and one of the tools for doing that is to make claims about safety; Anthropic models are safe and restricted to use by entities they deem safe, open models are not safe because anybody can use them and also who knows what those Chinese people are putting in their models.
And again, this also has nuance, models, including the Chinese open models, could be adversarial and we may not know it. Anthropic proved models can be a risk by sabotaging Fable briefly, causing it to produce bad results based on what the model thought it was being used for. This is why I tend to take Anthropic's words with a grain of salt. They're literally doing the unsafe things they say are risks of open models, while still laying claim to the "safe AI company" mantle.
Re: Muse Spark 1.3
#445llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle" https://tools.simonwillison.net/markdown-svg-renderer?url=ht... 4.2266 cents, 38 seconds. For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat. UPDATE: Here's an…
Would it not make more sense, assuming the purpose is to have a quick smoke test of model quality...to do a different animal, in a different setting each time, so as to defeat any tuning for your benchmark? Then go back and do the same for other models? Keep the pelican as a side baseline?
Re: Muse Spark 1.3
#446Earlier quoted context omitted.
At this point the only thing they're useful for is visualizing the differences between effort levels and roughly tracking the progression of models within a specific model family. And they still do that really well!
I don't see how useful this benchmark at all is for tracking the progression of models. I am not intending to bash on you personally but this is useless. People who are using AI models everyday are for sure not interested how close the AI model can visualize the pelican but they are interested in how they will perform on their daily tasks at work or private use. Correlation between doing good on pelican task and doin…
Re: Muse Spark 1.3
#447Ha, even with monitoring engineers keystrokes and mouse movements not SotA on OSWorld.
Re: Muse Spark 1.3
#448Earlier quoted context omitted.
+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks. (I'm not happy about the above being true, but it's the reality I seem…
And Fable 5.x makes Opus 5 look pretty dim, despite benchmarks suggesting they're comparable. The benchmarks really are just kinda meaningless.
Benchmarks are useful but only on a log2 basis. One model performing at 50% and another at 75% is just as impressive as one model performing at 78% and another at 90%. Confoundingly, a benchmark becomes useless once a frontier model scores over ~95% on them.
Re: Muse Spark 1.3
#449Earlier quoted context omitted.
I don't see how useful this benchmark at all is for tracking the progression of models. I am not intending to bash on you personally but this is useless. People who are using AI models everyday are for sure not interested how close the AI model can visualize the pelican but they are interested in how they will perform on their daily tasks at work or private use. Correlation between doing good on pelican task and doin…
Look at the difference between the Muse 1.2 and Muse 1.3 results.
Re: Muse Spark 1.3
#450Nice, Muse Spark is so good and keeps improving, but it's still not the best choice for any use-case. The Sol models are in their own league currently in terms of cost/speed/performance. Good improvements from 1.1 and 1.2[0], but when I tested 1.3 it was very slow (through openrouter). [0]: https://aibenchy.com/compare/meta-muse-spark-1-3-high/meta-m...
Cannot agree more with the Sol models. Everything else I try just seems "dumb".