Live data from Hacker News

Meta got caught gaming AI benchmarks

theverge.com

1–10 of 171 posts

Re: Meta got caught gaming AI benchmarks

#5
Is LMArena junk now?

I thought there was an aspect where you run two models on the same user-supplied query. Surely this can't be gamed?

> “optimized for conversationality”

I don't understand what that means - how it gives it an LMArena advantage.

Re: Meta got caught gaming AI benchmarks

#8
post #2

I believe this was designed to flatter the prompter more / be more ingratiating. Which is a worry if true (what it says about the people doing the comparing).

There's no end to the possible vectors of human manipulation with this "open-weight" black box.

Re: Meta got caught gaming AI benchmarks

#9
The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative.

This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off.

Did Zuck push for the release? I'm sure they knew it wasn't ready yet.

Re: Meta got caught gaming AI benchmarks

#10
post #7
post #6

Earlier quoted context omitted.

Not even first, OpenAI got caught a while back

Do you have a source for this? That's interesting (if true).

They got the dataset from Epoch AI for one of the benchmarks and pinky swore that they wouldn't train on it

https://techcrunch.com/2025/01/19/ai-benchmarking-organizati...

Post reply on HN