Live data from Hacker News

Meta got caught gaming AI benchmarks

theverge.com

51–60 of 171 posts

Re: Meta got caught gaming AI benchmarks

#51
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

"made an ambitious bet on MoEs"? No, DeepSeek is MoE, and they succeeded. Meta is not betting on MoE, it just does what other people have done.

Re: Meta got caught gaming AI benchmarks

#52

In other news, the head of AI research just left https://www.cnbc.com/2025/04/01/metas-head-of-ai-research-an...

I would have thought that title would belong to Yann.

TBH I'm very surprised Yann Le Cun is still there. He looks to me like a free thinker and an independent person. I don't think he buys into the Trump agenda and US nationalistic anti-Europe speech like Zuck does. He may be giving Zuck the benefit of the doubt, and probably is grateful that Zuck gave him a chance when nobody else did.

Re: Meta got caught gaming AI benchmarks

#53
post #25
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

I remember reading that they were in panic mode when the DeepSeek model came out so they must have scrambled and had to re-work a lot of things since DeepSeek was so competitive and open source as well

Fear of R2 looms large as well. I suspect they succumbed to the nuance collapse along the lines of “Is double checking results worth it if DeepSeek eats our lunch?”

Re: Meta got caught gaming AI benchmarks

#54
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

I don't know about Llama 4. Competition is intense in this field so you can't expect everybody to be number 1. However, I think the performance culture at Meta is counterproductive. Incentives are misaligned, I hope leadership will try to improve it. Employees are encouraged to ship half-baked features and move to another project. Quality isn't rewarded at all. The recent layoffs have made things even worse. Skilled…

I've never liked it, but

> Move fast and break things

is really a bad concept in this space, where you get limited shots at releasing something that generates interest.

> Employees are encouraged to ship half-baked features

And this is why I never liked that motto and have always pushed back at startups where I was hired that embraced this line of thought. Quality matters. It's context-dependent, so sometimes it matters a lot, and sometimes hardly. But "moving fast and breaking things" should be a deliberate choice, made for every feature, module, sprint, story all over again, IMO. If at all.

Re: Meta got caught gaming AI benchmarks

#55
post #40
post #27

Earlier quoted context omitted.

then why would they cheat?

they're all cheating, see grok

Are you referring to this [1]?

> Critics have pointed out that xAI’s approach involves running Grok 3 multiple times and cherry-picking the best output while comparing it against single runs of competitor models.

[1] https://medium.com/@cognidownunder/the-hype-machine-gpt-4-5-...

Re: Meta got caught gaming AI benchmarks

#56
post #34

I think it's most illustrative to see the sample battles (H2H) that LMArena released [1]. The outputs of Meta's model is too verbose and too 'yappy' IMO. And looking at the verdicts, it's no wonder by people are discounting LMArena rankings. [1]: https://huggingface.co/spaces/lmarena-ai/Llama-4-Maverick-03...

In fairness, 4o was like this until very recently. I suspect it comes from training on COT data from larger models.

Re: Meta got caught gaming AI benchmarks

#57
post #24
post #12

tech companies competing over something that is losing them money is the most bizarre spectacle yet.

I think Meta sees AI and VR/AR as a platform. They got left behind on the mobile platform and forever have to contend with Apple semi-monopoly. They have no control and little influence over the ecosystem. It's an existential threat to them. They have vowed not to make that mistake again so are pushing for an open future that won't be dominated by a few companies that could arbitrarily hurt Meta's business. That's th…

I'm in favor of whatever semi-monopoly enables fine grained permissions so Facebook can't en masse slurp Whatsapp (antitrust?) contacts

Re: Meta got caught gaming AI benchmarks

#58
post #54

Earlier quoted context omitted.

I don't know about Llama 4. Competition is intense in this field so you can't expect everybody to be number 1. However, I think the performance culture at Meta is counterproductive. Incentives are misaligned, I hope leadership will try to improve it. Employees are encouraged to ship half-baked features and move to another project. Quality isn't rewarded at all. The recent layoffs have made things even worse. Skilled…

I've never liked it, but > Move fast and break things is really a bad concept in this space, where you get limited shots at releasing something that generates interest. > Employees are encouraged to ship half-baked features And this is why I never liked that motto and have always pushed back at startups where I was hired that embraced this line of thought. Quality matters. It's context-dependent, so sometimes it matt…

I'd argue its a bad concept in any spaces that involve teams of people working together and deliverables that enter the real world.

Re: Meta got caught gaming AI benchmarks

#59
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

I mean, there's a reason they released it on a Saturday.
Post reply on HN