Live data from Hacker News

Meta got caught gaming AI benchmarks

theverge.com

21–30 of 171 posts

Re: Meta got caught gaming AI benchmarks

#21
post #10
post #7

Earlier quoted context omitted.

Do you have a source for this? That's interesting (if true).

They got the dataset from Epoch AI for one of the benchmarks and pinky swore that they wouldn't train on it https://techcrunch.com/2025/01/19/ai-benchmarking-organizati...

I don't see anything in the article about being caught. Maybe I missed something?

Re: Meta got caught gaming AI benchmarks

#22
post #14

Next on Matt Levine’s newsletter: Is Meta fudging with stock evaluation-correlated metrics? Is this securities fraud?

Sarcasm doesn't translate well in text. Please, elaborate.

Matt Levine has a common refrain which is that basically everything is securities fraud. If you were an investor who invested on the premise that Meta was good at AI and Zuck knowingly put out a bad model, is that securities fraud? Matt Levine will probably argue that it could be in a future edition of Money Stuff (his very good newsletter).

Re: Meta got caught gaming AI benchmarks

#23

Is LMArena junk now? I thought there was an aspect where you run two models on the same user-supplied query. Surely this can't be gamed? > “optimized for conversationality” I don't understand what that means - how it gives it an LMArena advantage.

It can be easily gamed. The users are self-selected, and they have zero incentive to be honest or rigorous or provide good responses. Some have incentives the opposite way. (There was a report of a prediction market user who said they had won a market on Gemini models by manipulating the votes; LMArena swore furiously there had definitely been no manipulation but was conspicuously silent on any details.) And the release of more LMArena responses has shown that a lot of the user ratings are blatantly wrong: either they're basically fraudulent, or LMArena's current users are people whose ratings you should be optimizing against because they are so ignorant, lazy, and superficial.

At this point, when I look at the output from my Gemini-2.5-pro sessions, they are so high quality, and take so long to read, and check, and have an informed opinion on, I just can't trust the slapdash approach of LMArena in assuming that careless driveby maybe-didn't-even-read-the-responses-ain't-no-one-got-time-for-that-nerd-shit ratings mean much of anything. There have been red flags in the past and I've been taking them ever less seriously even as one of many benchmarks since early last year, but maybe this is the biggest backfire yet. And it's only going to get worse. At this rate, without major overhaul, you should take being #1 on LMArena seriously as useful and important news - as a reason to not use a model.

It's past time for LMArena people to sit down and have some thorough reflection on whether it is still worth running at all, and at what point they are doing more harm than good. No benchmark lives forever and it is normal and healthy to shut them down at some point after having been saturated, but some manage to live a lot longer than they should have...

Re: Meta got caught gaming AI benchmarks

#24
post #12

tech companies competing over something that is losing them money is the most bizarre spectacle yet.

I think Meta sees AI and VR/AR as a platform. They got left behind on the mobile platform and forever have to contend with Apple semi-monopoly. They have no control and little influence over the ecosystem. It's an existential threat to them.

They have vowed not to make that mistake again so are pushing for an open future that won't be dominated by a few companies that could arbitrarily hurt Meta's business.

That's the stated rationale at least and I think it more or less makes sense

Re: Meta got caught gaming AI benchmarks

#25
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

I remember reading that they were in panic mode when the DeepSeek model came out so they must have scrambled and had to re-work a lot of things since DeepSeek was so competitive and open source as well

Re: Meta got caught gaming AI benchmarks

#27
post #12

tech companies competing over something that is losing them money is the most bizarre spectacle yet.

Come on, you can do the critical thinking here to understand why these companies would want the best in class (open/closed) weight LLMs.

then why would they cheat?

Re: Meta got caught gaming AI benchmarks

#28
post #22

Earlier quoted context omitted.

Sarcasm doesn't translate well in text. Please, elaborate.

Matt Levine has a common refrain which is that basically everything is securities fraud. If you were an investor who invested on the premise that Meta was good at AI and Zuck knowingly put out a bad model, is that securities fraud? Matt Levine will probably argue that it could be in a future edition of Money Stuff (his very good newsletter).

is it securities fraud? sort of.

If Mark, both through Meta and through his own resources, has the capital to hire and retain the best AI researchers / teams, and claims he's doing so, but puts out a model that sucks, he's liable. It's probably not directly fraud, but if he claims he's trying to compete with Google or Microsoft or Apple or whoever, yet doesn't adequately deploy a comparable amount of resources, capital, people, whatever, and doesn't explain why, it could (stretch) be securities fraud....I think.

Re: Meta got caught gaming AI benchmarks

#29

Is LMArena junk now? I thought there was an aspect where you run two models on the same user-supplied query. Surely this can't be gamed? > “optimized for conversationality” I don't understand what that means - how it gives it an LMArena advantage.

In one of karpathys videos he said that he was a bit suspicious that the models that score the highest in LMarena aren't the ones that people use the most to solve actual day to day problems.

Re: Meta got caught gaming AI benchmarks

#30
post #12

tech companies competing over something that is losing them money is the most bizarre spectacle yet.

Borderline conspiracy theory with an ounce of truth:

None of the models Meta put out are actually open source (by any measure), and everyone who are redistributing Llama models or any derivatives, or use Llama models for their business, are on the hook of getting sued in the future based on the terms and conditions people been explicitly/implicitly agreeing to when they use/redistribute these models.

If you start depending on these Llama models which have unfavorable proprietary terms today but Meta don't act on them, doesn't mean they won't act on it in the future. Maybe this has all been a play to get people into this position, so Meta can in the future start charging for them or something else.

Post reply on HN