Live data from Hacker News

Meta got caught gaming AI benchmarks

theverge.com

101–110 of 171 posts

Re: Meta got caught gaming AI benchmarks

#101
post #49
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

> it makes you wonder what they're going to do next They're just gonna keep throwing money at it. This is a hobby and talent magnet for them, instagram is the money printer. They've been working on VR for like a decade with barely much results in terms of users (compared to costs). This will be no different.

Both are also decent long-terms bets. Being VR market leader now means they will be VR market-leader with plenty of inhouse talent and IP when the technology matures and the market grows. Being in the AI race, even if they are not leading, means they have in-house talent and technology to be able to react to wherever the market is going with AI. They have one of the biggest messengers and one of the biggest image-posting sites, there is a decent chance AI will become important to them in some not-yet-obvious way.

One of Meta's biggest strengths is Zuckerberg being able to play these kinds of bets. Those bets being great for PR and talent acquisition is the cherry on top

Re: Meta got caught gaming AI benchmarks

#102
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

I don't know about Llama 4. Competition is intense in this field so you can't expect everybody to be number 1. However, I think the performance culture at Meta is counterproductive. Incentives are misaligned, I hope leadership will try to improve it. Employees are encouraged to ship half-baked features and move to another project. Quality isn't rewarded at all. The recent layoffs have made things even worse. Skilled…

> Employees are encouraged to ship half-baked features and move to another project

Maybe there is more to that. It's been more than a year since Llama 3 was released. That should be enough time for Meta to release something with significantly improvement. Or you mean quarter by quarter the engineers had to show that they were making impact in their perf review, which could be detrimental to the Llama 4 project?

Another thing that puzzles me is that again and again we see that the quality of a model can improve if we have more high-quality data, yet can't Meta manage to secure massive amount of new high-quality data to boost their model performance?

Re: Meta got caught gaming AI benchmarks

#103
post #7
post #6

Earlier quoted context omitted.

Not even first, OpenAI got caught a while back

Do you have a source for this? That's interesting (if true).

It happens with basically all papers on all topics. Benchmarks are useful when they are first introduced and used to measure things that were released before the benchmark. After that their usefulness rapidly declines.

Re: Meta got caught gaming AI benchmarks

#104
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

They knew they can't beat DeepSeek 2 months ago https://old.reddit.com/r/LocalLLaMA/comments/1i88g4y/meta_pa...

Re: Meta got caught gaming AI benchmarks

#105
post #81

Earlier quoted context omitted.

I don't know about Llama 4. Competition is intense in this field so you can't expect everybody to be number 1. However, I think the performance culture at Meta is counterproductive. Incentives are misaligned, I hope leadership will try to improve it. Employees are encouraged to ship half-baked features and move to another project. Quality isn't rewarded at all. The recent layoffs have made things even worse. Skilled…

For those who haven't heard of it, "The Hawthorne Effect" is the name given to a phenomena where when a person or group being studied is aware they are being studied, their performance goes up but as much as 50% for 4-8 weeks, then regresses to its norm. This is true if they are just being observed, or if some novel new processes are introduced. If the new things are beneficial, the performance rises for 4-8 weeks as…

I mean, it sounds like we should add in the McNamara fallacy also.

Re: Meta got caught gaming AI benchmarks

#106
post #76

Earlier quoted context omitted.

GPT 4o images is the future of all image gen. Every other player: Black Forest Labs' Flux, Stability.ai's Stable Diffusion, and even closed models like Ideogram and Midjourney, are all on the path to extinction. Image generation and editing must be multimodal. Full stop. Google Imagen will probably be the first model to match the capabilities of 4o. I'm hoping one of the open weights labs or Chinese AI giants will re…

One very important distinction between image models is the implementation: 4o is autogressive, slow, and extremely expensive. Although the Ghibli trend is market validation, I suspect that competitors may not want to copy it just yet.

Extremely expensive in what since? In that it costs $.03 instead of $.00003c? Yeah it's relatively far more expensive than other solutions, but from an absolute standpoint still very cheap for the vast majority of use cases. And it's a LOT better.

Re: Meta got caught gaming AI benchmarks

#107
post #49

Earlier quoted context omitted.

> it makes you wonder what they're going to do next They're just gonna keep throwing money at it. This is a hobby and talent magnet for them, instagram is the money printer. They've been working on VR for like a decade with barely much results in terms of users (compared to costs). This will be no different.

Both are also decent long-terms bets. Being VR market leader now means they will be VR market-leader with plenty of inhouse talent and IP when the technology matures and the market grows. Being in the AI race, even if they are not leading, means they have in-house talent and technology to be able to react to wherever the market is going with AI. They have one of the biggest messengers and one of the biggest image-pos…

This assumes no upstart will create a game changing innovation which upends everything.

Companies become complacent and confused when they get too big. Employees become trapped in a maze of performative careerism, and customer focus and a realistic understanding of threats from potential competitors both disappear.

It's been a consistent pattern since the earliest days of computing.

Nothing coming out of Big Tech at the moment is encouraging me to revise that assumption.

Re: Meta got caught gaming AI benchmarks

#109

Earlier quoted context omitted.

One very important distinction between image models is the implementation: 4o is autogressive, slow, and extremely expensive. Although the Ghibli trend is market validation, I suspect that competitors may not want to copy it just yet.

Extremely expensive in what since? In that it costs $.03 instead of $.00003c? Yeah it's relatively far more expensive than other solutions, but from an absolute standpoint still very cheap for the vast majority of use cases. And it's a LOT better.

Dall-E is already 4-8 cents per image. Afaik this is not in the API yet but I wouldn't be surprised if it's $1 or more.

Re: Meta got caught gaming AI benchmarks

#110

Earlier quoted context omitted.

I don't know about Llama 4. Competition is intense in this field so you can't expect everybody to be number 1. However, I think the performance culture at Meta is counterproductive. Incentives are misaligned, I hope leadership will try to improve it. Employees are encouraged to ship half-baked features and move to another project. Quality isn't rewarded at all. The recent layoffs have made things even worse. Skilled…

> Employees are encouraged to ship half-baked features and move to another project Maybe there is more to that. It's been more than a year since Llama 3 was released. That should be enough time for Meta to release something with significantly improvement. Or you mean quarter by quarter the engineers had to show that they were making impact in their perf review, which could be detrimental to the Llama 4 project? Anoth…

> Or you mean quarter by quarter the engineers had to show that they were making impact in their perf review

This is what I think they were referencing. Launching things looks nice in review packets and few to none are going to look into the quality of the output. Submitting your own self review means that you can cherry pick statistics and how you present them. That's why that culture incentivizes launching half baked products and moving on to something else because it's smart and profitable (launch yet another half baked project) to distance yourself from the half baked project you started.

Post reply on HN