Lot more details in the linked report https://ai.meta.com/static-resource/muse-spark-1-1-evaluatio... From Terminal-bench-2.1 details, > We use a bash-tool-only agent harness to evaluate 89 Terminal-Bench 2.1 tasks from the official repository, where resources are capped at 6 CPU cores and 8GB RAM. This disqualifies the results. Each terminal bench task has a cpu upper limit and RAM upper limit. Overriding either is…
Muse Spark 1.1
121–130 of 228 posts
Re: Muse Spark 1.1
#122How is every company able to show itself at the top of every benchmark?
Re: Muse Spark 1.1
#123My trust factor is gone with Meta right now. Has there been any independent analysis to confirm they didn't cheat on benchmarks again?
They cheated again: https://news.ycombinator.com/item?id=48847019
Re: Muse Spark 1.1
#124Earlier quoted context omitted.
> At least in China a lot of software developers are now struggling. Do you think that Chinese software industry is that relevant to the kind of software market talked about on HN? I.e. lots of enterprise b2b and infra companies. Chinese companies have always had a very low willingness to pay for software which kinda breaks the flywheel of B2B SaaS companies and companies to service those companies all the way down.
Its a signal. They were earning well and AI crashed the market in China.
They have had real issues with deflation rather than the inflation most Western countries have seen over the past five years.
Re: Muse Spark 1.1
#125The pricing is insane: $1.25/$4.5 for 1M tokens, and $0.15 for cached input! https://dev.meta.ai/docs/getting-started/pricing-rate-limits
Re: Muse Spark 1.1
#126Lot more details in the linked report https://ai.meta.com/static-resource/muse-spark-1-1-evaluatio... From Terminal-bench-2.1 details, > We use a bash-tool-only agent harness to evaluate 89 Terminal-Bench 2.1 tasks from the official repository, where resources are capped at 6 CPU cores and 8GB RAM. This disqualifies the results. Each terminal bench task has a cpu upper limit and RAM upper limit. Overriding either is…
Huh? What are you talking about? https://www.anthropic.com/engineering/infrastructure-noise Is anthropic benchmark maxxing and cheating on terminal bench too? They don't follow the strict resource "limits" either
Re: Muse Spark 1.1
#127I personally do not like Meta, but I'll say this. The more competition, the better for regular consumers. (Enterprise too) - Chinese models - Grok - Meta - Google - OpenAI - Anthropic I think this is a win. I'm building like crazy to take advantage of all these subsidized tokens while I can.
Yeah, I think it is definitely great. Having said that, I am still debating in my mind whether the volume of software engineers needed in the AI era is going to increase or decrease because of all of these advancements. On the one hand, because it is easy to build products, more and more people will build. And more and more products and features will be built. However, a lot of people who are non-technical will also…
What I really want is Claude as a deep part of the operating system.
If that happens then a whole lot of the abstraction of software vanishes along with what we think of today as software jobs. I think many new forms of knowledge work would emerge from this though.
I would think that needs massive local compute but I can't imagine that is not the future down the line.
Re: Muse Spark 1.1
#128Earlier quoted context omitted.
Huh? What are you talking about? https://www.anthropic.com/engineering/infrastructure-noise Is anthropic benchmark maxxing and cheating on terminal bench too? They don't follow the strict resource "limits" either
What that link describes is basically the motivation to go from terminal bench 2.0 to 2.1. The latter simply fixed the common issues/complaints. There is a long github discussion on tbench's about it
Sure for old tasks you could argue that now its not required to boost because infra errors are alleviated with better default limits. My point more so is that its a strange thing to index on because if you wanted to cheat on the benchmark, it does not particularly seem like something that shifts results? Once the API is out maybe I'll eat my words, but I don't really believe that if you manually tried to reproduce the results with lower limits you'd see significantly different results
Re: Muse Spark 1.1
#129The pricing is insane: $1.25/$4.5 for 1M tokens, and $0.15 for cached input! https://dev.meta.ai/docs/getting-started/pricing-rate-limits
This is still ridiculously expensive imagine having to pay $10 for 100 search results on Google, thats essentially what this is. I really dont see how anyone's willing spend more than $1.50 per mm output. Let alone $15-50. Does anyone actually pay for usage based billing as a consumer?
https://platform.claude.com/docs/en/about-claude/pricing
Model Base Input Tokens 5m Cache Writes 1h Cache Writes Cache Hits & Refreshes Output Tokens
Claude Fable 5 $10 / MTok $12.50 / MTok $20 / MTok $1 / MTok $50 / MTok
Claude Opus 4.8 $5 / MTok $6.25 / MTok $10 / MTok $0.50 / MTok $25 / MTok
Note Fable costs $50 MTok and Opus 4.8 costs $25 / MTok.
Re: Muse Spark 1.1
#130Earlier quoted context omitted.
They cheated again: https://news.ycombinator.com/item?id=48847019
No they didn't. Please read: https://www.anthropic.com/engineering/infrastructure-noise
> The extra resources enable the agent to try approaches that only work with generous allocations, such as pulling in large dependencies, spawning expensive subprocesses, and running memory-intensive test suites.
> An agent that writes lean, efficient code very fast will do well under tight constraints. An agent that brute-forces solutions with heavyweight tools will do well under generous ones. Both are legitimate things to test, but collapsing them into a single score without specifying the resource configuration makes the differences—and real-world generalizability—hard to interpret.
So changing the resource limits changes the benchmark. Yet their score table claims their score to be for Terminal-Bench 2.1, not Terminal-Bench 2.1 with raised limits.