Live data from Hacker News

Muse Spark 1.1

ai.meta.com

111–120 of 228 posts

Re: Muse Spark 1.1

#112
post #72

Interesting how the prevalent opinion until yesterday seems to have been that OpenAI & Anthropic are irreversibly ahead and now with xAI and Meta at least delivered something that's competitive with useful models and cheap too. Granted, the narrative that the two leading labs are ahead still holds with Fable (and perhaps an upcoming GPT6), but it's not as over as common knowledge by the opinion leaders would have us…

People misinterpreted Google being behind as Anthropic and OpenAi being really ahead, when it was really just Google falling behind the same way it did with Tensorflow, Angular and GCP.

> when it was really just Google falling behind the same way it did with Tensorflow, Angular and GCP

Not sure I agree. Angular fell behind in popularity but was (is? unsure atm) still eminently usable. I gave gemini a test drive recently and it was horrendous, as in "picking dirt cheap Chinese model over gemini any day" bad, and with overzealous guardrails to boot. 3.1 pro feels a year behind and is extremely lazy. 3.5 flash feels like a model you’d run on your 128gb macbook, not something that was released a month ago and which costs a fair bit when used through api.

In any case: as of right now I think that we went from a three horse race to anthropic / openai as premium choices vs whatever is the Chinese fotm for a fraction of the cost. 3.5 pro better be a miracle if google wants to hang out with the big boys, otherwise their only strategy is hoping that both US labs go broke and they remain the last man standing.

Re: Muse Spark 1.1

#113
post #45
post #28

Earlier quoted context omitted.

To expand on Chinese models: - DeepSeek - GLM (Z.ai) - Minimax - Kimi (Moonshot) - Hy3 (Tencent) - Qwen (Alibaba) (Each one of these with weights available to download and run locally)

GLM 5.2 is great, but is so rate limited now I no longer recommend it

Rumors are Nvidia H200s got approved so infrastructure might be improving soon.

Re: Muse Spark 1.1

#116
post #50

How is every company able to show itself at the top of every benchmark?

First look what models are worse in a set of self selected benchmarks.

Second, compare to older versions of competitor s models.

Still does not look good? Compare to own previous models.

Re: Muse Spark 1.1

#117

Earlier quoted context omitted.

Weren't they caught multiple times gaming the benchmark even more so then the rest?

Let me assure you, literally everybody does this

I don't think it even matters. Because noone will continue to use an LLM that doesn't work well for them, whether or not it has a good bench result. So for their own sake, the correct representation can actually win them some loyalty:

eg. Model X is weaker than Fable, but competes well with Opus/Sonnet and costs 1/5th as much etc - something similar playing out with Grok 4.5.

Re: Muse Spark 1.1

#118

> Model API is not available in your region. :( Well, Vietnam is not in the list of restricted territories. Anyway, what is "your region" ? Is this where I am now, or is it where I activated my Oculus 2 five years ago ?

Same in Argentina. It's almost surely a region whitelist for now (it's the only reason Argentina ever gets blocked).

Re: Muse Spark 1.1

#120

Lot more details in the linked report https://ai.meta.com/static-resource/muse-spark-1-1-evaluatio... From Terminal-bench-2.1 details, > We use a bash-tool-only agent harness to evaluate 89 Terminal-Bench 2.1 tasks from the official repository, where resources are capped at 6 CPU cores and 8GB RAM. This disqualifies the results. Each terminal bench task has a cpu upper limit and RAM upper limit. Overriding either is…

Huh? What are you talking about?

https://www.anthropic.com/engineering/infrastructure-noise

Is anthropic benchmark maxxing and cheating on terminal bench too? They don't follow the strict resource "limits" either

Post reply on HN