Live data from Hacker News

Meta got caught gaming AI benchmarks

theverge.com

31–40 of 171 posts

Re: Meta got caught gaming AI benchmarks

#31

Is LMArena junk now? I thought there was an aspect where you run two models on the same user-supplied query. Surely this can't be gamed? > “optimized for conversationality” I don't understand what that means - how it gives it an LMArena advantage.

There are almost certainly ways to fine-tune the model in ways that make it perform better on the Arena, but perform worse in other benchmarks or in practice. Usually that's not a good trade-off. What's being suggested here is that Meta is running such a fine-tuned version on the Arena (and reporting those numbers) while running models with different fine-tuning on other benchmarks (and reporting those numbers), while giving the appearance that those are actually the same models.

Re: Meta got caught gaming AI benchmarks

#33
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

Do you know that they made a bet on MoE? Meaning they abandonded dense models? I doubt that is the case. Just releasing MoE Llama 4 does not constitute a "bet" without further information.

Also from what I can tell this performs better than models with parameter counts equal to one expert, and worse than fully dense models equal to total parameter count. Isn't that kind of what we'd expect? in what way is that a failure?

Maybe I am missing some details. But it feels like you have an axe to grind.

Re: Meta got caught gaming AI benchmarks

#34
I think it's most illustrative to see the sample battles (H2H) that LMArena released [1]. The outputs of Meta's model is too verbose and too 'yappy' IMO. And looking at the verdicts, it's no wonder by people are discounting LMArena rankings.

[1]: https://huggingface.co/spaces/lmarena-ai/Llama-4-Maverick-03...

Re: Meta got caught gaming AI benchmarks

#36
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

Do you know that they made a bet on MoE? Meaning they abandonded dense models? I doubt that is the case. Just releasing MoE Llama 4 does not constitute a "bet" without further information. Also from what I can tell this performs better than models with parameter counts equal to one expert, and worse than fully dense models equal to total parameter count. Isn't that kind of what we'd expect? in what way is that a fail…

A 4x8 MOE performs better than an 8B but worse than a 32B, is your statement?

My response would be, "so why bother with MOE?"

However deepseek r1 is MOE from my understanding, but the "E" are all =>32B parameters. There's > 20 experts. I could be misinformed; however, even so, I'd say a MOE with 32B or even 70B experts will outperform (define this!) Models with equal parameter counts, because deepseek outperforms (define?) ChatGPT et al.

Re: Meta got caught gaming AI benchmarks

#37
post #28
post #22

Earlier quoted context omitted.

Matt Levine has a common refrain which is that basically everything is securities fraud. If you were an investor who invested on the premise that Meta was good at AI and Zuck knowingly put out a bad model, is that securities fraud? Matt Levine will probably argue that it could be in a future edition of Money Stuff (his very good newsletter).

is it securities fraud? sort of. If Mark, both through Meta and through his own resources, has the capital to hire and retain the best AI researchers / teams, and claims he's doing so, but puts out a model that sucks, he's liable. It's probably not directly fraud, but if he claims he's trying to compete with Google or Microsoft or Apple or whoever, yet doesn't adequately deploy a comparable amount of resources, capit…

And the fine for that? Probably 0.001% of revenue. If that.

Re: Meta got caught gaming AI benchmarks

#38
post #27

Earlier quoted context omitted.

Come on, you can do the critical thinking here to understand why these companies would want the best in class (open/closed) weight LLMs.

then why would they cheat?

Well we'll see if they suffer consequences of this and they cheated too hard, but being perceived as best in class is arguably worth even more than being the best in class, especially if differences in performance are hard to perceive anecdotally.

The goal is long term control over a technology's marketshare, as winner take all dynamics are in play here.

Post reply on HN