Live data from Hacker News

Muse Code and Muse Spark 1.2

research.meta.ai

251–260 of 265 posts

Re: Muse Code and Muse Spark 1.2

#251
post #155

Earlier quoted context omitted.

If you look at papers on benchmarks, they're usually created to expose gaps in how models are trained. It should be no surprise that models get better on them over time, because you can't get better at what you don't measure. Cherry picking the benchmarks you present is where the falsehoods lie.

Another thing that sort of puzzles me about benchmarks is that LLMs are not deterministic and do not always complete a problem. So what are the results actually representing? The best run? The average? It is all in some ways a falsehood

That's a good point and conventionally if benchmarks aren't run as "one shot", it is denoted as "benchmark@K". Inference time scaling has historically shown improvement.

Generally though, many of these fairness complaints do go away if there is "3rd party testing". Right now, companies reporting their own benchmarks has all the problems that 3rd party testing resolves in many other industries.

Re: Muse Code and Muse Spark 1.2

#252
post #248

Earlier quoted context omitted.

does the contract enable the customer to monitor/search Meta to ensure they are honoring the contract? if there is no mechanism for that it means very little. Though I bet/hope some will feed them "watermarked"/unique but worthless things and watch for traces of that to pop up in models or something like that, but that's hatdly enough to just take their word for it.

That tends to be what discovery in lawsuits is for.

That's circular reasoning, how would there be a lawsuit if customer have no way of knowing?

Re: Muse Code and Muse Spark 1.2

#253

Earlier quoted context omitted.

"they 'trust me'. dumb fucks." The only difference between students doing it then and professionals doing it now is the students had no positive, glaring reason to mistrust.

I can't understand the psychology of people who think like this. You really think a throwaway quote when Zuckerberg was a college student applies nowadays? You really think Meta would risk getting a massive (losing) lawsuit on their hands, in exchange for what? Middling amounts of training data?

> You really think a throwaway quote when Zuckerberg was a college student applies nowadays?

What does "nowadays" mean? What changed?

> You really think Meta would risk getting a massive (losing) lawsuit on their hands, in exchange for what?

And what is the risk? Food companies are at risk when they put addictive chemicals into food because they get controlled regularly. What is the equivalent here?

Re: Muse Code and Muse Spark 1.2

#254
post #170
post #169

Earlier quoted context omitted.

Sounds like a great way to get your account revoked.

How come?

Because it'll get declined after a few attempts once it's past the limit, then your account would likely get marked for fraud. Try it and report back, but repeated method of payment failures get accounts suspended pretty reliably at any place that does CC processing, since it's also a risk to their merchant account to have a higher percentage of declined payments (ie higher fraud risk) and it's far safer for them to cut you off / blacklist you vs giving you more chances and the sure bet that they'll invoke higher merchant account / payment processor fees. If they make it to the "high risk" category, they're cooked too; Mastercard for example has MATCH that once you're added, you're considered high-risk at ANY other place you try to get CC processing going again. You can read a bit more about that here: https://docs.stripe.com/disputes/match Payment processors can fully drop merchants that have too high of a risk / too many chargebacks / too high payment failure percentages.

Re: Muse Code and Muse Spark 1.2

#256
post #93

Earlier quoted context omitted.

I've been surprised by the reception to this, as OpenAI, for a while now, has had free API usage when data sharing is enabled ( https://help.openai.com/en/articles/10306912-sharing-feedbac... )

I enabled all data sharing settings but still don’t have a message about free use on that screen - the help page says free tokens are available to “some” users - is that 1% of users, 40% of users, etc? Does your screen have the message that you’re getting free tokens? https://platform.openai.com/settings/organization/data-contr...

I had no idea about this before. I just enabled it.

  You're eligible for free daily usage on traffic shared with OpenAI.

      Up to 250 thousand tokens per day across gpt-5.4, gpt-5.2, gpt-5.1, gpt-5.1-codex, gpt-5, gpt-5-codex, gpt-5-chat-latest, gpt-4.1, gpt-4o, o1, and o3
      Up to 2.5 million tokens per day across gpt-5.4-mini, gpt-5.4-nano, gpt-5.1-codex-mini, gpt-5-mini, gpt-5-nano, gpt-4.1-mini, gpt-4.1-nano, gpt-4o-mini, o1-mini, o3-mini, o4-mini, and codex-mini-latest.

  Usage beyond these limits, as well as usage for other models, will be billed at standard rates. Some limitations apply. Learn more.

Re: Muse Code and Muse Spark 1.2

#257

I registered on their platform API, and immediately it asked for a selfie inage to verify Im a human. Thanks, but there is definately no chance I gave them my selfie to use their model.

you generally should use intermediaries (they buy API from the main provider and serve it to you) when dealing with a company like Meta, a company you want no direct relationship with.

they have the spark 1.2 "contributor" which lets them train on your data which is cheaper than deepseek flash atm. I always assume they keep my data no matter what they say so it's a deal.

Re: Muse Code and Muse Spark 1.2

#258

Earlier quoted context omitted.

No one has to say it and plan it out loud. But the data will be sitting there. The incentive to improve the model for enterprise use will only get stronger. It doesn't take much for one engineer or team to go rogue to hack a benchmark. There was a whole cheating controversy with llama 4.

What does cheating benchmarks have to do with breaching financial user agreements? To use your argument: all it takes is one whistleblower to get the company sued for billions of dollars.

What does unethical behavior have to do with unethical behavior? What does looking at data you're not supposed to look at have to do with looking at data you're not supposed look at? Please don't be obtuse.

People whistleblow on Meta all the time. Most recently they got fined half a billion for suppressing child safety research. Kids getting groomed. It doesn't make a difference to a company that has a net profit of $60 billion a year.

Re: Muse Code and Muse Spark 1.2

#260

Earlier quoted context omitted.

"they 'trust me'. dumb fucks." The only difference between students doing it then and professionals doing it now is the students had no positive, glaring reason to mistrust.

I can't understand the psychology of people who think like this. You really think a throwaway quote when Zuckerberg was a college student applies nowadays? You really think Meta would risk getting a massive (losing) lawsuit on their hands, in exchange for what? Middling amounts of training data?

Meta literally ran a man in the middle attack to spy on it's user's encrypted network traffic when they used third party apps [0]. More recently (and relevant to this issue), the engaged in industrial scale piracy to get training data for their LLMs [1]. The idea that they have changed since Zuck was a college student creeping on his female classmates and now wouldn't commit actual crimes against their own users in order to get a bit more data is just demonstrably false.

[0] https://www.techradar.com/computing/cyber-security/facebooks...

[1] https://www.tomshardware.com/tech-industry/artificial-intell...

Post reply on HN