Live data from Hacker News

Muse Spark: Scaling towards personal superintelligence

ai.meta.com

371–380 of 392 posts

Re: Muse Spark: Scaling towards personal superintelligence

#371
post #318

Comes impressively close to GPT 5.4 / Gemini 3.1 Pro / Opus 4.6! Mostly behind OpenAI on coding/agentic benchmarks, behind Google on text reasoning, behind Anthropic on Humanity's Last Exam with tools (surprisingly the only benchmark where Anthropic leads currently). Meta hasn’t fully caught up, but they came close and I think can solidly claim to be a frontier lab again. I’d call it a 3.5 horse race right now, and h…

Grok code was my daily driver for months while it was free and it was fantastic - it is certainly no worse than it was a few months ago. Unfortunately with LLMs everything is based off your use case, domain and the context you give it. I also use Grok daily for health questions as the other models are too afraid to give input on medical matters

Why do you need to ask any AI questions regarding your health every day?

Re: Muse Spark: Scaling towards personal superintelligence

#373
post #355

Earlier quoted context omitted.

[flagged]

If you are trying to come up with anti-media conspiracies there are always plenty of ways to do it against any media company. The idea that NY Times is particularly anti-Meta seems a stretch. They - like most traditional media companies - are anti-tech in general. The fact they also collect data doesn't make their reporting untrue. Personally I think a much more interesting rumor to make up would be that Yann Lecun (…

>They - like most traditional media companies - are anti-tech in general. The fact they also collect data doesn't make their reporting untrue.

(sigh) In olden times you would have been free to use the em dash as you pleased. Unfortunately, now it's considered signal that you're an AI bot.

Re: Muse Spark: Scaling towards personal superintelligence

#374
post #364
post #203

Earlier quoted context omitted.

Anthropic generally seem more into living within market discipline and market signals of some sort. Products with margins, even if it's sort of irrelevant considering R&D costs and capital inflow. That said, there's nothing like the real thing. The risk is something like the railroad bubble and the dotcom. Over-investement, circular revenue and a timeline that doesn't work. Or, maybe it'll work out.

The weird position they find themselves in now is that they have to keep making it smarter... but they already made it too smart (Mythos). I'm not sure how that's going to work out exactly. They find an arbitrary intelligence cutoff point between Opus and Mythos, label it "acceptable risk", and then the labs coordinate to gradually nudge that line forward and hope the internet doesn't break?

> but they already made it too smart (Mythos).

It's largely a marketing tactic. It will be released, and it won't be long before other models show similar capabilities.

If they wanted they could add guardrails. The scales required to brute force search for vulnerabilities like they did would be very identifiable.

Re: Muse Spark: Scaling towards personal superintelligence

#375
post #310

Earlier quoted context omitted.

They really weren't horrible. They were ~gpt4o, with the added benefit that you could run them on premise. Just "regular" models, non "thinking". Inefficient architecture (number of active out of total) but otherwise "decent" models. They got trashed online by bots and chinese shills (I was online that weekend when it happened, it's something to behold). Just because they were non-thinking when thinking was clearly t…

> They were ~gpt4o, with the added benefit that you could run them on premise. No, they are bad models. They were benchmaxxed on LMAreana and a few other benchmarks but as soon as you try them yourself they fall to pieces. I have my own agentic benchmark[1] I use to compare models. Llama-4-scout-17b-16e scores 14/25, while llama-4-maverick-17b-128e scores 12/25. By comparison gemma-4-E4B-it-GGUF:Q4_K_M scores 15/25 (…

> By comparison gemma-4-E4B-it-GGUF:Q4_K_M scores 15/25 (that is a 4B parameter model!)

Gemma 4 E4B is slightly confusingly named, its a 8B param model

Re: Muse Spark: Scaling towards personal superintelligence

#376

> Muse Spark is a natively multimodal reasoning model with support for [...] visual chain of thought [...]. Do they mean "the chain of thought is visible to the user" (ie. not hidden like ChatGPT), or "the medium of the chain of thought is not text, but visuals" (ie. thinking in images). I'd guess the former, since it wouldn't be economical to generate transient images, just for thinking. But I'm not sure why they'd…

Perhaps more importantly, will their chain of thought be "real"? So far the ones I've seen seem to be elaborate fakery. They look good unless you dig in at which point you often find that it merely looks plausible on the surface but that something else is going on under the hood.

I don't know what you mean by that. We know what's going on under the hood always: linear algebra, the attention mechanism etc.

To my first approximation all "Chain of thought" means is that instead of having to prompt the model to discuss everything in text and then decide at the end[1], now it sort of automatically does that so you don't need to prompt it.

[1] Which used to bring about very substantial improvements in performance on some tasks

Re: Muse Spark: Scaling towards personal superintelligence

#377
post #364

Earlier quoted context omitted.

The weird position they find themselves in now is that they have to keep making it smarter... but they already made it too smart (Mythos). I'm not sure how that's going to work out exactly. They find an arbitrary intelligence cutoff point between Opus and Mythos, label it "acceptable risk", and then the labs coordinate to gradually nudge that line forward and hope the internet doesn't break?

> but they already made it too smart (Mythos). It's largely a marketing tactic. It will be released, and it won't be long before other models show similar capabilities. If they wanted they could add guardrails. The scales required to brute force search for vulnerabilities like they did would be very identifiable.

Scam Altman already pulled this trick numerous times.

Whats wrong with people? Is it really that hard to see the truth?

Re: Muse Spark: Scaling towards personal superintelligence

#378
post #203

Earlier quoted context omitted.

I suspect this is the real reason behind Anthropic limiting subscriptions to their own products and keeping API prices several times higher than comparable models. Applications more sticky than API users and less technical users more sticky than programmers (ie Cowork more sticky than Code).

Anthropic generally seem more into living within market discipline and market signals of some sort. Products with margins, even if it's sort of irrelevant considering R&D costs and capital inflow. That said, there's nothing like the real thing. The risk is something like the railroad bubble and the dotcom. Over-investement, circular revenue and a timeline that doesn't work. Or, maybe it'll work out.

The whole premise is based on the fact that over-investing in GPUs and models are a good thing here as it yields more 'intelligence'.

This as it turned out was not true for rail roads - more and more rail roads isnt a good thing.

The real dilemma facing the model producers is that all this money invested for a general model, targeting general intelligence, is a disaster and essentially the investment into existing assets is a write off. Then on top of that if this is true, youve got data centres full of compute that aren't being used up.

Re: Muse Spark: Scaling towards personal superintelligence

#379
Do we have any numbers on input, output and conversation context window limit?

I tried multiple riddles, graphs and questions I know some LLMs fails at, but this one seems to do well. But I still don't have much trust in Meta after the scandal of them fiddling with their previous models to look good.

Re: Muse Spark: Scaling towards personal superintelligence

#380

Earlier quoted context omitted.

Regardless they are getting that revenue through genuine demand for their product. It’s not like they are selling back some commodity product, billions are being spent on model outputs. I think anyone who has used Opus 4.6 can see what is causing this demand. It is genuinely “smart” in the sense that it can work its way around non-trivial coding problems.

But at some point even if the product is useful if it costs twice what is getting in, won’t that be a problem ?

I don't see why tokens/$ would suddenly stop dropping. Maybe this is the first time the cost of compute will plateau, but do have any reason to think so?
Post reply on HN