Live data from Hacker News

Meta got caught gaming AI benchmarks

theverge.com

111–120 of 171 posts

Re: Meta got caught gaming AI benchmarks

#112

Earlier quoted context omitted.

A 4x8 MOE performs better than an 8B but worse than a 32B, is your statement? My response would be, "so why bother with MOE?" However deepseek r1 is MOE from my understanding, but the "E" are all =>32B parameters. There's > 20 experts. I could be misinformed; however, even so, I'd say a MOE with 32B or even 70B experts will outperform (define this!) Models with equal parameter counts, because deepseek outperforms (de…

Easy, vastly improved inference performance on machines with larger RAM but lower bandwidth/compute. These are becoming more popular such as Apple's M series chips, AMD's strix halo series, and the upcoming DGX Spark from Nvidia.

yes i understand all that. I was saying the claim is incorrect. My understanding of deepseek is mechanically correct but apparently they use 3B models as experts, per your sibling comment. I don't buy it, regardless of what they put in the paper - 3B models are pretty dumb, and R1 isn't dumb. No amount of shuffling between "dumb" experts will make the output not dumb. it's more likely 32x32B experts, based on the quant sizes i've seen.

A deepseek employee is welcome to correct me.

Re: Meta got caught gaming AI benchmarks

#114
post #10

Earlier quoted context omitted.

They got the dataset from Epoch AI for one of the benchmarks and pinky swore that they wouldn't train on it https://techcrunch.com/2025/01/19/ai-benchmarking-organizati...

I don't see anything in the article about being caught. Maybe I missed something?

[flagged]

Re: Meta got caught gaming AI benchmarks

#116

Is LMArena junk now? I thought there was an aspect where you run two models on the same user-supplied query. Surely this can't be gamed? > “optimized for conversationality” I don't understand what that means - how it gives it an LMArena advantage.

LMArena was always junk. I work in this space and while the media takes it seriously most scientists don't. Random people ask random stuff and then it measures how good they feel. This is only a worthwhile evaluation if you're Google or Meta or OpenAI and you need to make a chartbot that keeps people coming back. It doesn't measure anything else useful.

Llama 1 derived models on it were beating gpt 3.5 by have less refusals.

Re: Meta got caught gaming AI benchmarks

#117

Earlier quoted context omitted.

I don't know about Llama 4. Competition is intense in this field so you can't expect everybody to be number 1. However, I think the performance culture at Meta is counterproductive. Incentives are misaligned, I hope leadership will try to improve it. Employees are encouraged to ship half-baked features and move to another project. Quality isn't rewarded at all. The recent layoffs have made things even worse. Skilled…

I agree. I think of it like a car engine. You can push it up to a certain RPM and it will keep making more and more power. Above that RPM, the engine starts to produce less power and eventually blows a gasket. I think the performance-based management worked for a while because there were some gains to be had by pushing people harder. However, they’ve gone past that and are now pushing people too hard and getting wors…

The problem is, there's always some engine pushing the power envelope, or a person pushing their performance harder. And then the rest of them have to keep up.

Re: Meta got caught gaming AI benchmarks

#118

Earlier quoted context omitted.

I don't see anything in the article about being caught. Maybe I missed something?

[flagged]

>They got the dataset from Epoch AI for one of the benchmarks and pinky swore that they wouldn't train on it

Is anything here actually false or do you not like the conclusions that people may draw from it?

Re: Meta got caught gaming AI benchmarks

#119
post #110

Earlier quoted context omitted.

> Employees are encouraged to ship half-baked features and move to another project Maybe there is more to that. It's been more than a year since Llama 3 was released. That should be enough time for Meta to release something with significantly improvement. Or you mean quarter by quarter the engineers had to show that they were making impact in their perf review, which could be detrimental to the Llama 4 project? Anoth…

> Or you mean quarter by quarter the engineers had to show that they were making impact in their perf review This is what I think they were referencing. Launching things looks nice in review packets and few to none are going to look into the quality of the output. Submitting your own self review means that you can cherry pick statistics and how you present them. That's why that culture incentivizes launching half bak…

I like how Netflix set up its incentive systems years ago. Essentially they told the employees that all they needed to do is deliver what the company wanted. It was perfectly okay that an employee did their job and didn't move up or do more. Per their chief talent officer McCord, "a manager's job is all about setting the context" and the employees were let loose to deliver. This method puts a really high bar on the managers, as the entire report chain must know clearly what they want delivered. Their expectation must be high enough to move the company forward, but not too ridiculous to turn Netflix into a burnout factory.
Post reply on HN