Earlier quoted context omitted.
"I'm not bitter! No chip on my shoulder."
bitter about what? I'm a long time employee
Meta got caught gaming AI benchmarks
111–120 of 171 posts
Re: Meta got caught gaming AI benchmarks
#112Earlier quoted context omitted.
A 4x8 MOE performs better than an 8B but worse than a 32B, is your statement? My response would be, "so why bother with MOE?" However deepseek r1 is MOE from my understanding, but the "E" are all =>32B parameters. There's > 20 experts. I could be misinformed; however, even so, I'd say a MOE with 32B or even 70B experts will outperform (define this!) Models with equal parameter counts, because deepseek outperforms (de…
Easy, vastly improved inference performance on machines with larger RAM but lower bandwidth/compute. These are becoming more popular such as Apple's M series chips, AMD's strix halo series, and the upcoming DGX Spark from Nvidia.
A deepseek employee is welcome to correct me.
Re: Meta got caught gaming AI benchmarks
#113tech companies competing over something that is losing them money is the most bizarre spectacle yet.
Re: Meta got caught gaming AI benchmarks
#114Earlier quoted context omitted.
They got the dataset from Epoch AI for one of the benchmarks and pinky swore that they wouldn't train on it https://techcrunch.com/2025/01/19/ai-benchmarking-organizati...
I don't see anything in the article about being caught. Maybe I missed something?
Re: Meta got caught gaming AI benchmarks
#115Re: Meta got caught gaming AI benchmarks
#116Is LMArena junk now? I thought there was an aspect where you run two models on the same user-supplied query. Surely this can't be gamed? > “optimized for conversationality” I don't understand what that means - how it gives it an LMArena advantage.
LMArena was always junk. I work in this space and while the media takes it seriously most scientists don't. Random people ask random stuff and then it measures how good they feel. This is only a worthwhile evaluation if you're Google or Meta or OpenAI and you need to make a chartbot that keeps people coming back. It doesn't measure anything else useful.
Re: Meta got caught gaming AI benchmarks
#117Earlier quoted context omitted.
I don't know about Llama 4. Competition is intense in this field so you can't expect everybody to be number 1. However, I think the performance culture at Meta is counterproductive. Incentives are misaligned, I hope leadership will try to improve it. Employees are encouraged to ship half-baked features and move to another project. Quality isn't rewarded at all. The recent layoffs have made things even worse. Skilled…
I agree. I think of it like a car engine. You can push it up to a certain RPM and it will keep making more and more power. Above that RPM, the engine starts to produce less power and eventually blows a gasket. I think the performance-based management worked for a while because there were some gains to be had by pushing people harder. However, they’ve gone past that and are now pushing people too hard and getting wors…
Re: Meta got caught gaming AI benchmarks
#118Earlier quoted context omitted.
I don't see anything in the article about being caught. Maybe I missed something?
[flagged]
Is anything here actually false or do you not like the conclusions that people may draw from it?
Re: Meta got caught gaming AI benchmarks
#119Earlier quoted context omitted.
> Employees are encouraged to ship half-baked features and move to another project Maybe there is more to that. It's been more than a year since Llama 3 was released. That should be enough time for Meta to release something with significantly improvement. Or you mean quarter by quarter the engineers had to show that they were making impact in their perf review, which could be detrimental to the Llama 4 project? Anoth…
> Or you mean quarter by quarter the engineers had to show that they were making impact in their perf review This is what I think they were referencing. Launching things looks nice in review packets and few to none are going to look into the quality of the output. Submitting your own self review means that you can cherry pick statistics and how you present them. That's why that culture incentivizes launching half bak…
Re: Meta got caught gaming AI benchmarks
#120In other news, the head of AI research just left https://www.cnbc.com/2025/04/01/metas-head-of-ai-research-an...