Live data from Hacker News

Muse Spark 1.1

ai.meta.com

71–80 of 228 posts

Re: Muse Spark 1.1

#71

Earlier quoted context omitted.

> but they can't read minds I don't know who needs to hear this, but neither can humans. You've implicitly assumed here that AI systems will always be worse at contextualizing and framing questions than the average engineer. I'm not sure that it's even true now let alone in the nebulous future. You haven't narrowed the fundamental myopia of the assumption here, just dressed it in slightly different clothing.

> You've implicitly assumed here that AI systems will always be worse at contextualizing and framing questions than the average engineer. How would they know what to ask or contextualize if they don't know what the user wants?

Are you suggesting that psychic mindreading powers are real?

> How would they know

How would you? The answer is the same.

Re: Muse Spark 1.1

#72
Interesting how the prevalent opinion until yesterday seems to have been that OpenAI & Anthropic are irreversibly ahead and now with xAI and Meta at least delivered something that's competitive with useful models and cheap too. Granted, the narrative that the two leading labs are ahead still holds with Fable (and perhaps an upcoming GPT6), but it's not as over as common knowledge by the opinion leaders would have us believe.

Re: Muse Spark 1.1

#73
post #69

Lot more details in the linked report https://ai.meta.com/static-resource/muse-spark-1-1-evaluatio... From Terminal-bench-2.1 details, > We use a bash-tool-only agent harness to evaluate 89 Terminal-Bench 2.1 tasks from the official repository, where resources are capped at 6 CPU cores and 8GB RAM. This disqualifies the results. Each terminal bench task has a cpu upper limit and RAM upper limit. Overriding either is…

Out of curiosity, how often are the resource limits the bottlenecks? What do harnesses do to help here - limit parallelism? More efficient tools?

The task could be verifiable in the environment so limiting its CPU and RAM could be to discourage brute forcing the answer.

Re: Muse Spark 1.1

#74

Earlier quoted context omitted.

> You've implicitly assumed here that AI systems will always be worse at contextualizing and framing questions than the average engineer. How would they know what to ask or contextualize if they don't know what the user wants?

Are you suggesting that psychic mindreading powers are real? > How would they know How would you? The answer is the same.

I don't understand what you mean. I can't build software I can't describe.

If you're implying chatbots can ask their "client" what to build, good luck with that—contractors are at least liable for what they produce and have extreme incentives to ensure that their clients are happy. To the extent of refusing to build anything if they don't know what they want....

Re: Muse Spark 1.1

#75
post #45
post #28

Earlier quoted context omitted.

To expand on Chinese models: - DeepSeek - GLM (Z.ai) - Minimax - Kimi (Moonshot) - Hy3 (Tencent) - Qwen (Alibaba) (Each one of these with weights available to download and run locally)

GLM 5.2 is great, but is so rate limited now I no longer recommend it

Aren't there multiple providers for it? is it rate limited in all providers?

Re: Muse Spark 1.1

#76

Everyone has been loving to shit on the Alexander Wang acquisition but this seems legitimately impressive to me? Meta's AI org when from a total mismanaged dumpster fire for multiple years to delivering a competitive model in less than a year on essentially their first try?

Not their first try. There’s been reporting about how they’ve kept pushing their model releases back because of underwhelming performance.

Re: Muse Spark 1.1

#77

Lot more details in the linked report https://ai.meta.com/static-resource/muse-spark-1-1-evaluatio... From Terminal-bench-2.1 details, > We use a bash-tool-only agent harness to evaluate 89 Terminal-Bench 2.1 tasks from the official repository, where resources are capped at 6 CPU cores and 8GB RAM. This disqualifies the results. Each terminal bench task has a cpu upper limit and RAM upper limit. Overriding either is…

Why are resource limits considered at all aside from models accidentally fork bombing themselves?

I thought the benchmark was supposed to be about terminal use and specifically chaining together lots of bash tool calls. Which test cases does this matter for?

Re: Muse Spark 1.1

#78
post #31

I personally do not like Meta, but I'll say this. The more competition, the better for regular consumers. (Enterprise too) - Chinese models - Grok - Meta - Google - OpenAI - Anthropic I think this is a win. I'm building like crazy to take advantage of all these subsidized tokens while I can.

While data centers are still using lots of energy created from fossil fuels and many still evaporate water for cooling? No wonder we still can’t get climate change under control

> No wonder we still can’t get climate change under control

This is was historically a money issue, being green used to be wildly more expensive.

Now being green is cheaper, the limiting factor is how fast PV and batteries can be made or imported.

Recent reports of the sum of all US data centres currently in planning, has a power demand exceeding the (capacity-factor-adjusted!) global annual supply of new PV.

This would be less of a problem, but still a problem, if Trump wasn't trying to get in the way of anything green, or if the companies building data centres decided to also support factories to make more PV.

* Planned new demand: 300 GW; PV factory capacity ~ 600 GW nameplate, but the capacity factor is 14% so that's really 84 GW on average.

Re: Muse Spark 1.1

#79
post #53

Everyone has been loving to shit on the Alexander Wang acquisition but this seems legitimately impressive to me? Meta's AI org when from a total mismanaged dumpster fire for multiple years to delivering a competitive model in less than a year on essentially their first try?

How is it their first try? They were leading the race with Llama 3.x a few years ago.

They were leading the race in a niche category a few years ago. Now they are, according to some benchmarks, even on the right playing field.

Re: Muse Spark 1.1

#80
post #64

How are people trying this? I don't see it on openrouter. Any ways of testing this without subscribing to meta stuff?

Probably need to wait some hours/1-2 days and openrouter will add it.

Thanks. I was asking because I couldn't find even their previous 1.0 model there.
Post reply on HN