Earlier quoted context omitted.
LMArena was always junk. I work in this space and while the media takes it seriously most scientists don't. Random people ask random stuff and then it measures how good they feel. This is only a worthwhile evaluation if you're Google or Meta or OpenAI and you need to make a chartbot that keeps people coming back. It doesn't measure anything else useful.
I hear AI news from time to time from the M5M in the US - and the only place I've ever seen "LMArena" is on HN and in the LM studio discord. At a ratio of 5:1 at least.
Meta got caught gaming AI benchmarks
131–140 of 171 posts
Re: Meta got caught gaming AI benchmarks
#132Earlier quoted context omitted.
I've never liked it, but > Move fast and break things is really a bad concept in this space, where you get limited shots at releasing something that generates interest. > Employees are encouraged to ship half-baked features And this is why I never liked that motto and have always pushed back at startups where I was hired that embraced this line of thought. Quality matters. It's context-dependent, so sometimes it matt…
> is really a bad concept in this space, where you get limited shots at releasing something that generates interest. It's a really bad concept in any space. We would be living in a better world if Zuck had, at least once, thought "Maybe we shouldn't do that".
I struggle with this because it feels like so many 'rules' in the world where the important half remains unsaid. That unsaid portion is then mediated by goodharts law.
If the other half is 'then slow down and learn something' its really not that bad, nothing is sacred, we try we fail we learn we (critical) don't repeat the mistake. Thats human learning - we don't learn from mistakes we learn from reflecting on mistakes.
But if learning isn't part of the loop - if its a self justifying defense for fuckups, if the unsaid remains unsaid, its a disaster waiting to happen.
The difference is usually in what you reward. If you reward ship, you get the defensive version - and you will ship crap. If you reward institutional knowledge building you don't. Engineers are often taught that 'good, fast, or cheap pick 2'. The reality is its usually closer to 1 or 1.5. If you pick fast...you get fast.
Re: Meta got caught gaming AI benchmarks
#133Earlier quoted context omitted.
> Or you mean quarter by quarter the engineers had to show that they were making impact in their perf review This is what I think they were referencing. Launching things looks nice in review packets and few to none are going to look into the quality of the output. Submitting your own self review means that you can cherry pick statistics and how you present them. That's why that culture incentivizes launching half bak…
I like how Netflix set up its incentive systems years ago. Essentially they told the employees that all they needed to do is deliver what the company wanted. It was perfectly okay that an employee did their job and didn't move up or do more. Per their chief talent officer McCord, "a manager's job is all about setting the context" and the employees were let loose to deliver. This method puts a really high bar on the m…
> employees that all they needed to do is deliver what the company wanted
How did this work out in practice and across teams? My experience at Meta within my team was that it would be almost impossible to determine what the company actually wanted from our team in a year. Goals kept changing and the existing incentive system works against this since other teams are trying to come up with their own solutions to things which may impact your team.
> an employee did their job and didn't move up
Does Netflix cull employees if they haven't reached a certain IC level? I know at Meta SWEs need to reach IC5 after a while or risk being culled.
Re: Meta got caught gaming AI benchmarks
#134Earlier quoted context omitted.
Extremely expensive in what since? In that it costs $.03 instead of $.00003c? Yeah it's relatively far more expensive than other solutions, but from an absolute standpoint still very cheap for the vast majority of use cases. And it's a LOT better.
Dall-E is already 4-8 cents per image. Afaik this is not in the API yet but I wouldn't be surprised if it's $1 or more.
Re: Meta got caught gaming AI benchmarks
#135In other news, the head of AI research just left https://www.cnbc.com/2025/04/01/metas-head-of-ai-research-an...
I would have thought that title would belong to Yann.
Re: Meta got caught gaming AI benchmarks
#136Earlier quoted context omitted.
Tangent, but does anyone know why links to lmarena.ai are banned on Reddit, site-wide? (Last I checked 1 month ago)
evidence for this assertion please? its highly unlikely and maybe a bug
It was a really strange situation that lasted for months. Posts or comments with a direct link were removed like they were the worst of the worst.
I tried posting to r/bugs and it was downvoted immediately. I eventually contacted he lmarena folks, so maybe they resolved it with reddit.
Re: Meta got caught gaming AI benchmarks
#137Re: Meta got caught gaming AI benchmarks
#138Meta got caught _first_.
According to the article, Meta publicly stated, right below the benchmark comparison, that the version of Llama on LMArena was the experimental chat version:
> According to Meta’s own materials, it deployed an “experimental chat version” of Maverick to LMArena that was specifically “optimized for conversationality”
The AI benchmark in question, LMArena, compares Llama 4 experimental to closed models like ChatGPT 4o latest, and Llama performs better (https://lmarena.ai/?leaderboard).
Re: Meta got caught gaming AI benchmarks
#139Re: Meta got caught gaming AI benchmarks
#140Earlier quoted context omitted.
I don't see anything in the article about being caught. Maybe I missed something?
[flagged]
> If an MMLU question asks about an abstract algebra proof, is it cheating to have trained on papers about abstract algebra?
This kind of disingenuous bullshit is exactly why people call you cheaters.
> Generally, I don’t think anyone here is cheating and I think we’re relatively diligent with our evals.
You guys should follow Apple's cult guidelines: never stand out. Think different