How is every company able to show itself at the top of every benchmark?
Muse Spark 1.1
111–120 of 228 posts
Re: Muse Spark 1.1
#112Interesting how the prevalent opinion until yesterday seems to have been that OpenAI & Anthropic are irreversibly ahead and now with xAI and Meta at least delivered something that's competitive with useful models and cheap too. Granted, the narrative that the two leading labs are ahead still holds with Fable (and perhaps an upcoming GPT6), but it's not as over as common knowledge by the opinion leaders would have us…
People misinterpreted Google being behind as Anthropic and OpenAi being really ahead, when it was really just Google falling behind the same way it did with Tensorflow, Angular and GCP.
Not sure I agree. Angular fell behind in popularity but was (is? unsure atm) still eminently usable. I gave gemini a test drive recently and it was horrendous, as in "picking dirt cheap Chinese model over gemini any day" bad, and with overzealous guardrails to boot. 3.1 pro feels a year behind and is extremely lazy. 3.5 flash feels like a model you’d run on your 128gb macbook, not something that was released a month ago and which costs a fair bit when used through api.
In any case: as of right now I think that we went from a three horse race to anthropic / openai as premium choices vs whatever is the Chinese fotm for a fraction of the cost. 3.5 pro better be a miracle if google wants to hang out with the big boys, otherwise their only strategy is hoping that both US labs go broke and they remain the last man standing.
Re: Muse Spark 1.1
#113Earlier quoted context omitted.
To expand on Chinese models: - DeepSeek - GLM (Z.ai) - Minimax - Kimi (Moonshot) - Hy3 (Tencent) - Qwen (Alibaba) (Each one of these with weights available to download and run locally)
GLM 5.2 is great, but is so rate limited now I no longer recommend it
Re: Muse Spark 1.1
#114Re: Muse Spark 1.1
#115Re: Muse Spark 1.1
#116How is every company able to show itself at the top of every benchmark?
Second, compare to older versions of competitor s models.
Still does not look good? Compare to own previous models.
Re: Muse Spark 1.1
#117Earlier quoted context omitted.
Weren't they caught multiple times gaming the benchmark even more so then the rest?
Let me assure you, literally everybody does this
eg. Model X is weaker than Fable, but competes well with Opus/Sonnet and costs 1/5th as much etc - something similar playing out with Grok 4.5.
Re: Muse Spark 1.1
#118> Model API is not available in your region. :( Well, Vietnam is not in the list of restricted territories. Anyway, what is "your region" ? Is this where I am now, or is it where I activated my Oculus 2 five years ago ?
Re: Muse Spark 1.1
#119My trust factor is gone with Meta right now. Has there been any independent analysis to confirm they didn't cheat on benchmarks again?
Re: Muse Spark 1.1
#120Lot more details in the linked report https://ai.meta.com/static-resource/muse-spark-1-1-evaluatio... From Terminal-bench-2.1 details, > We use a bash-tool-only agent harness to evaluate 89 Terminal-Bench 2.1 tasks from the official repository, where resources are capped at 6 CPU cores and 8GB RAM. This disqualifies the results. Each terminal bench task has a cpu upper limit and RAM upper limit. Overriding either is…
https://www.anthropic.com/engineering/infrastructure-noise
Is anthropic benchmark maxxing and cheating on terminal bench too? They don't follow the strict resource "limits" either