Meta got caught gaming AI benchmarks
121–130 of 171 posts
Re: Meta got caught gaming AI benchmarks
#122Earlier quoted context omitted.
I would have thought that title would belong to Yann.
TBH I'm very surprised Yann Le Cun is still there. He looks to me like a free thinker and an independent person. I don't think he buys into the Trump agenda and US nationalistic anti-Europe speech like Zuck does. He may be giving Zuck the benefit of the doubt, and probably is grateful that Zuck gave him a chance when nobody else did.
Zuck doesn't buy it, either. He just knows what's good for business right now.
In an example of the worst person you know making a great point, Josh Hawley said "What really struck me is that they can read an election return." [0].
Though it's worth remembering, it's very difficult to accumulate the volume of data necessary to do the current kind of AI training while sticking to the strictest interpretations of EU privacy law. Social media companies aren't just feeding the user data into marketing algorithms, they're feeding them into AI models. If you're a leading researcher in that field - Like Le Cun - and the current state-of-the-art means getting as much data as possible, you might not appreciate the regulatory environment of the EU.
[0] https://www.npr.org/2025/02/27/nx-s1-5302712/senator-josh-ha...
Re: Meta got caught gaming AI benchmarks
#123The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…
I don't know about Llama 4. Competition is intense in this field so you can't expect everybody to be number 1. However, I think the performance culture at Meta is counterproductive. Incentives are misaligned, I hope leadership will try to improve it. Employees are encouraged to ship half-baked features and move to another project. Quality isn't rewarded at all. The recent layoffs have made things even worse. Skilled…
You're not encouraged per se to ship half-baked features, but if you don't have enough "impact" at the end of the half (for mid cycle checkin) or year (for full PSC cycle) then you're going to get "Below Expectations" and then "Meets Most" (or worse) and with the current environment a swift offboarding.
When I was there (working in integrity) our group of staff+ engineers opined how it led to perverse incentives - and whilst you can work there and do great work, and get good ratings, I saw too many examples of "optimizing for PSC" (otherwise known as PSC hacking).
Re: Meta got caught gaming AI benchmarks
#124I'm just shocked that the companies who stole all kinds of copyrighted material would again do something unethical to keep the bubble and gravy train going...
Imagine this but you remove the noise and can walk like in an art gallery (it's a diffusion model but LLMs can be loosely converted into 3D maps with objects, too): https://writings.stephenwolfram.com/2023/07/generative-ai-sp...
Re: Meta got caught gaming AI benchmarks
#125Earlier quoted context omitted.
I don't know about Llama 4. Competition is intense in this field so you can't expect everybody to be number 1. However, I think the performance culture at Meta is counterproductive. Incentives are misaligned, I hope leadership will try to improve it. Employees are encouraged to ship half-baked features and move to another project. Quality isn't rewarded at all. The recent layoffs have made things even worse. Skilled…
I've never liked it, but > Move fast and break things is really a bad concept in this space, where you get limited shots at releasing something that generates interest. > Employees are encouraged to ship half-baked features And this is why I never liked that motto and have always pushed back at startups where I was hired that embraced this line of thought. Quality matters. It's context-dependent, so sometimes it matt…
It's a really bad concept in any space.
We would be living in a better world if Zuck had, at least once, thought "Maybe we shouldn't do that".
Re: Meta got caught gaming AI benchmarks
#126Is LMArena junk now? I thought there was an aspect where you run two models on the same user-supplied query. Surely this can't be gamed? > “optimized for conversationality” I don't understand what that means - how it gives it an LMArena advantage.
(Last I checked 1 month ago)
Re: Meta got caught gaming AI benchmarks
#127The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…
It's not a big deal. Llama 4 feels like a flop because the expectations are really high based on their previous releases and the sense of momentum in the ecosystem because of DeepSeek. At the end of the day, LLama 4 didn't meet the elevated expectations, but they're fine. They'll continue to improve and iterate and maybe the next one will be more hype worthy, or maybe expectations will be readjusted as the specter of…
Re: Meta got caught gaming AI benchmarks
#128Earlier quoted context omitted.
> They cant produce new things, and so meta and google are struggling to compete when actual merit matters, and they cant just brute force the solutions. You haven't been keeping up. Less than 2 weeks ago, Google released a model that has crushed the competition, clearly being SotA while currently effectively free for personal use. Gemini 2.0 was already good, people just weren't paying attention. In fact 1.5 pro was…
gemini 2.5 pro isnt good, and if you think it is, you arent using LLMs correctly. The model gets crushed by o1 pro and sonnet 3.7 thinking. Build a large contextual prompt ( > 50k tokens) with a ton of code, and see how bad it is. I cancelled my gemini subscription
Re: Meta got caught gaming AI benchmarks
#129Is LMArena junk now? I thought there was an aspect where you run two models on the same user-supplied query. Surely this can't be gamed? > “optimized for conversationality” I don't understand what that means - how it gives it an LMArena advantage.
Tangent, but does anyone know why links to lmarena.ai are banned on Reddit, site-wide? (Last I checked 1 month ago)
Re: Meta got caught gaming AI benchmarks
#130tech companies competing over something that is losing them money is the most bizarre spectacle yet.
I think Meta sees AI and VR/AR as a platform. They got left behind on the mobile platform and forever have to contend with Apple semi-monopoly. They have no control and little influence over the ecosystem. It's an existential threat to them. They have vowed not to make that mistake again so are pushing for an open future that won't be dominated by a few companies that could arbitrarily hurt Meta's business. That's th…