Live data from Hacker News

Meta got caught gaming AI benchmarks

theverge.com

131–140 of 171 posts

Re: Meta got caught gaming AI benchmarks

#131

Earlier quoted context omitted.

LMArena was always junk. I work in this space and while the media takes it seriously most scientists don't. Random people ask random stuff and then it measures how good they feel. This is only a worthwhile evaluation if you're Google or Meta or OpenAI and you need to make a chartbot that keeps people coming back. It doesn't measure anything else useful.

I hear AI news from time to time from the M5M in the US - and the only place I've ever seen "LMArena" is on HN and in the LM studio discord. At a ratio of 5:1 at least.

It's mentioned quite a bit in the LLM related subreddits.

Re: Meta got caught gaming AI benchmarks

#132
post #54

Earlier quoted context omitted.

I've never liked it, but > Move fast and break things is really a bad concept in this space, where you get limited shots at releasing something that generates interest. > Employees are encouraged to ship half-baked features And this is why I never liked that motto and have always pushed back at startups where I was hired that embraced this line of thought. Quality matters. It's context-dependent, so sometimes it matt…

> is really a bad concept in this space, where you get limited shots at releasing something that generates interest. It's a really bad concept in any space. We would be living in a better world if Zuck had, at least once, thought "Maybe we shouldn't do that".

>It's a really bad concept in any space.

I struggle with this because it feels like so many 'rules' in the world where the important half remains unsaid. That unsaid portion is then mediated by goodharts law.

If the other half is 'then slow down and learn something' its really not that bad, nothing is sacred, we try we fail we learn we (critical) don't repeat the mistake. Thats human learning - we don't learn from mistakes we learn from reflecting on mistakes.

But if learning isn't part of the loop - if its a self justifying defense for fuckups, if the unsaid remains unsaid, its a disaster waiting to happen.

The difference is usually in what you reward. If you reward ship, you get the defensive version - and you will ship crap. If you reward institutional knowledge building you don't. Engineers are often taught that 'good, fast, or cheap pick 2'. The reality is its usually closer to 1 or 1.5. If you pick fast...you get fast.

Re: Meta got caught gaming AI benchmarks

#133
post #110

Earlier quoted context omitted.

> Or you mean quarter by quarter the engineers had to show that they were making impact in their perf review This is what I think they were referencing. Launching things looks nice in review packets and few to none are going to look into the quality of the output. Submitting your own self review means that you can cherry pick statistics and how you present them. That's why that culture incentivizes launching half bak…

I like how Netflix set up its incentive systems years ago. Essentially they told the employees that all they needed to do is deliver what the company wanted. It was perfectly okay that an employee did their job and didn't move up or do more. Per their chief talent officer McCord, "a manager's job is all about setting the context" and the employees were let loose to deliver. This method puts a really high bar on the m…

Unfortunately I wasn't able to get an interview with Netflix.

> employees that all they needed to do is deliver what the company wanted

How did this work out in practice and across teams? My experience at Meta within my team was that it would be almost impossible to determine what the company actually wanted from our team in a year. Goals kept changing and the existing incentive system works against this since other teams are trying to come up with their own solutions to things which may impact your team.

> an employee did their job and didn't move up

Does Netflix cull employees if they haven't reached a certain IC level? I know at Meta SWEs need to reach IC5 after a while or risk being culled.

Re: Meta got caught gaming AI benchmarks

#134

Earlier quoted context omitted.

Extremely expensive in what since? In that it costs $.03 instead of $.00003c? Yeah it's relatively far more expensive than other solutions, but from an absolute standpoint still very cheap for the vast majority of use cases. And it's a LOT better.

Dall-E is already 4-8 cents per image. Afaik this is not in the API yet but I wouldn't be surprised if it's $1 or more.

[deleted]

Re: Meta got caught gaming AI benchmarks

#135

In other news, the head of AI research just left https://www.cnbc.com/2025/04/01/metas-head-of-ai-research-an...

I would have thought that title would belong to Yann.

It's a misnomer - the VP left. Yann is the Chief Scientist, which I imagine most would agree would be the 'head' of a research division.

Re: Meta got caught gaming AI benchmarks

#136
post #129

Earlier quoted context omitted.

Tangent, but does anyone know why links to lmarena.ai are banned on Reddit, site-wide? (Last I checked 1 month ago)

evidence for this assertion please? its highly unlikely and maybe a bug

I should have re-tested prior to posting. It is fixed now. I just tried posting a comment with the link and it was not [removed by reddit].

It was a really strange situation that lasted for months. Posts or comments with a direct link were removed like they were the worst of the worst.

I tried posting to r/bugs and it was downvoted immediately. I eventually contacted he lmarena folks, so maybe they resolved it with reddit.

Re: Meta got caught gaming AI benchmarks

#138
post #4

Meta got caught _first_.

“Got caught” is a misleading way to present what happened.

According to the article, Meta publicly stated, right below the benchmark comparison, that the version of Llama on LMArena was the experimental chat version:

> According to Meta’s own materials, it deployed an “experimental chat version” of Maverick to LMArena that was specifically “optimized for conversationality”

The AI benchmark in question, LMArena, compares Llama 4 experimental to closed models like ChatGPT 4o latest, and Llama performs better (https://lmarena.ai/?leaderboard).

Re: Meta got caught gaming AI benchmarks

#139
For me at least, the 10M context window is a big deal and as long as it's decent, I'm going to use it instead. I'm running Scout locally and my chat history can get very long. I'm very frustrated when the context window runs out. I haven't been able to fully test the context length but at least that one isn't fudged.

Re: Meta got caught gaming AI benchmarks

#140

Earlier quoted context omitted.

I don't see anything in the article about being caught. Maybe I missed something?

[flagged]

Why are you being disingenuous? Simply having access to the eval in question is already enough for your synthetics guys to match the distribution, and of course you don't contaminate directly on train, that would be stupid, and you would get caught, but if it does inform the reward, the result is the same. You _should_ quit, but you wouldn't because you'd already convinced yourself you're doing RL God's work, not sleight of hand.

> If an MMLU question asks about an abstract algebra proof, is it cheating to have trained on papers about abstract algebra?

This kind of disingenuous bullshit is exactly why people call you cheaters.

> Generally, I don’t think anyone here is cheating and I think we’re relatively diligent with our evals.

You guys should follow Apple's cult guidelines: never stand out. Think different

Post reply on HN