Live data from Hacker News

Meta got caught gaming AI benchmarks

theverge.com

81–90 of 171 posts

Re: Meta got caught gaming AI benchmarks

#81
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

I don't know about Llama 4. Competition is intense in this field so you can't expect everybody to be number 1. However, I think the performance culture at Meta is counterproductive. Incentives are misaligned, I hope leadership will try to improve it. Employees are encouraged to ship half-baked features and move to another project. Quality isn't rewarded at all. The recent layoffs have made things even worse. Skilled…

For those who haven't heard of it, "The Hawthorne Effect" is the name given to a phenomena where when a person or group being studied is aware they are being studied, their performance goes up but as much as 50% for 4-8 weeks, then regresses to its norm.

This is true if they are just being observed, or if some novel new processes are introduced. If the new things are beneficial, the performance rises for 4-8 weeks as usual, but when it regresses it regresses to a higher performance reflecting the value of the new process.

But when poor management introduce a counter-productive change, the Hawthorne Effect makes it look like a resounding success for 4-8 weeks. Then the effect fades, and performance drops below the original level. Sufficiently devious managers either move on to new projects or blame the workers for failing to maintain the new higher pace of performance.

This explains a lot of the incentive for certain types of leaders to champion arbitrary changes, take a victory lap, and then disassociate themselves from accountability for the long-term success or failure of their initiative.

(There is quite a bit of controversy over what the mechanisms for the Hawthorne Effect are, and whether change alone can introduce it for whether participants need to feel they are being observed, but the model as I see it fits my anecdotal experience where new processes are always accompanied by attempts to meet new performance goals, and everyone is extremely aware that the outcome is being measured.)

Re: Meta got caught gaming AI benchmarks

#83
post #27

Earlier quoted context omitted.

Come on, you can do the critical thinking here to understand why these companies would want the best in class (open/closed) weight LLMs.

then why would they cheat?

I didn't see evidence of cheating in the article. Having a slightly differently tuned version of 4 is not the most dastardly thing that can be done. Everything else is insinuation.

Re: Meta got caught gaming AI benchmarks

#87
post #76

I tried to make Studio Ghibli inspired images using presumably their new models. It was ass.

GPT 4o images is the future of all image gen. Every other player: Black Forest Labs' Flux, Stability.ai's Stable Diffusion, and even closed models like Ideogram and Midjourney, are all on the path to extinction. Image generation and editing must be multimodal. Full stop. Google Imagen will probably be the first model to match the capabilities of 4o. I'm hoping one of the open weights labs or Chinese AI giants will re…

One very important distinction between image models is the implementation: 4o is autogressive, slow, and extremely expensive.

Although the Ghibli trend is market validation, I suspect that competitors may not want to copy it just yet.

Re: Meta got caught gaming AI benchmarks

#88
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

It's not a big deal. Llama 4 feels like a flop because the expectations are really high based on their previous releases and the sense of momentum in the ecosystem because of DeepSeek. At the end of the day, LLama 4 didn't meet the elevated expectations, but they're fine. They'll continue to improve and iterate and maybe the next one will be more hype worthy, or maybe expectations will be readjusted as the specter of…

The switching costs are so low (zero) that anyone using these models just jumps to the best performer. I also agree that this is not a brand or narrative sensitive project.

Re: Meta got caught gaming AI benchmarks

#89
post #6
post #4

Meta got caught _first_.

Not even first, OpenAI got caught a while back

People have been gaming ML benchmarks as long as there have been ML benchmarks. That's why it's better to see if other researchers are incorporating a technique into their actual models rather than 'is this paper the bold entry in a benchmark table'. But it takes longer.

Re: Meta got caught gaming AI benchmarks

#90
post #70

Earlier quoted context omitted.

> They cant produce new things, and so meta and google are struggling to compete when actual merit matters, and they cant just brute force the solutions. You haven't been keeping up. Less than 2 weeks ago, Google released a model that has crushed the competition, clearly being SotA while currently effectively free for personal use. Gemini 2.0 was already good, people just weren't paying attention. In fact 1.5 pro was…

gemini 2.5 pro isnt good, and if you think it is, you arent using LLMs correctly. The model gets crushed by o1 pro and sonnet 3.7 thinking. Build a large contextual prompt ( > 50k tokens) with a ton of code, and see how bad it is. I cancelled my gemini subscription

https://aider.chat/docs/leaderboards/ your experience doesn't align with my experience or this benchmark. o1 pro is good but I would rather do 20 cycles on gemini 2.5 rather than wait for Pro to return.
Post reply on HN