Live data from Hacker News

Meta got caught gaming AI benchmarks

theverge.com

61–70 of 171 posts

Re: Meta got caught gaming AI benchmarks

#62
post #54

Earlier quoted context omitted.

I don't know about Llama 4. Competition is intense in this field so you can't expect everybody to be number 1. However, I think the performance culture at Meta is counterproductive. Incentives are misaligned, I hope leadership will try to improve it. Employees are encouraged to ship half-baked features and move to another project. Quality isn't rewarded at all. The recent layoffs have made things even worse. Skilled…

I've never liked it, but > Move fast and break things is really a bad concept in this space, where you get limited shots at releasing something that generates interest. > Employees are encouraged to ship half-baked features And this is why I never liked that motto and have always pushed back at startups where I was hired that embraced this line of thought. Quality matters. It's context-dependent, so sometimes it matt…

> is really a bad concept in this space, where you get limited shots at releasing something that generates interest.

Sure, but long term effects are more depending on the actual performance of the model, than anything.

Say they launch a model that is hyped to be the best, but when people try it, it's worse than other models. People will quickly forget about it, unless it's particularly good at something.

Alternatively, say they launch a model that doesn't even get a press release, or any benchmark results published ahead of launch, but the model actually rocks at a bunch of use cases. People will start using it regardless of the initial release, and continue to do so as long as it's a best model.

Re: Meta got caught gaming AI benchmarks

#63
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

It's not a big deal. Llama 4 feels like a flop because the expectations are really high based on their previous releases and the sense of momentum in the ecosystem because of DeepSeek. At the end of the day, LLama 4 didn't meet the elevated expectations, but they're fine. They'll continue to improve and iterate and maybe the next one will be more hype worthy, or maybe expectations will be readjusted as the specter of diminishing returns continues to creep in.

Re: Meta got caught gaming AI benchmarks

#64
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

I don't know about Llama 4. Competition is intense in this field so you can't expect everybody to be number 1. However, I think the performance culture at Meta is counterproductive. Incentives are misaligned, I hope leadership will try to improve it. Employees are encouraged to ship half-baked features and move to another project. Quality isn't rewarded at all. The recent layoffs have made things even worse. Skilled…

It's also terrible output, even before you consider what looks like catastrophic forgetting from crappy RL. The emoji use and writing style make me want to suck-start a revolver. I don't know how they expect anyone to actually use it.

Re: Meta got caught gaming AI benchmarks

#65
post #30
post #12

tech companies competing over something that is losing them money is the most bizarre spectacle yet.

Borderline conspiracy theory with an ounce of truth: None of the models Meta put out are actually open source (by any measure), and everyone who are redistributing Llama models or any derivatives, or use Llama models for their business, are on the hook of getting sued in the future based on the terms and conditions people been explicitly/implicitly agreeing to when they use/redistribute these models. If you start dep…

You never go full Oracle.

Re: Meta got caught gaming AI benchmarks

#66
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

I don't know about Llama 4. Competition is intense in this field so you can't expect everybody to be number 1. However, I think the performance culture at Meta is counterproductive. Incentives are misaligned, I hope leadership will try to improve it. Employees are encouraged to ship half-baked features and move to another project. Quality isn't rewarded at all. The recent layoffs have made things even worse. Skilled…

I agree. I think of it like a car engine. You can push it up to a certain RPM and it will keep making more and more power. Above that RPM, the engine starts to produce less power and eventually blows a gasket.

I think the performance-based management worked for a while because there were some gains to be had by pushing people harder. However, they’ve gone past that and are now pushing people too hard and getting worse results. Every machine has its operating limits and an area where it operates most efficiently. A company is no different.

Re: Meta got caught gaming AI benchmarks

#67
The truth is that the vast majority of FAANG engineers making high six figures are only good at deterministic work. They cant produce new things, and so meta and google are struggling to compete when actual merit matters, and they cant just brute force the solutions. Inside these companies, the massive tech systems built, are actually generally terrible, but they pile on legions of engineers to fix the problems.

This is the culture of META hurting them, they are paying "AI VPs" millions of dollars to go to status meetings to get dates for when these models will be done. Meanwhile, deepseek r1 has a flat hierarchy with engineers that actually understand low level computing

Its making a mockery of big tech, and is why startups exist. Big company employees rise the ranks by building skill sets other than producing true economic value

Re: Meta got caught gaming AI benchmarks

#68

Earlier quoted context omitted.

I don't know about Llama 4. Competition is intense in this field so you can't expect everybody to be number 1. However, I think the performance culture at Meta is counterproductive. Incentives are misaligned, I hope leadership will try to improve it. Employees are encouraged to ship half-baked features and move to another project. Quality isn't rewarded at all. The recent layoffs have made things even worse. Skilled…

I agree. I think of it like a car engine. You can push it up to a certain RPM and it will keep making more and more power. Above that RPM, the engine starts to produce less power and eventually blows a gasket. I think the performance-based management worked for a while because there were some gains to be had by pushing people harder. However, they’ve gone past that and are now pushing people too hard and getting wors…

very nice analogy!

Re: Meta got caught gaming AI benchmarks

#69

Ahmad al-Dahle, who leads "Gen AI" at Meta, wrote this on Twitter: ... We're also hearing some reports of mixed quality across different services ... We've also heard claims that we trained on test sets -- that's simply not true and we would never do that. Our best understanding is that the variable quality people are seeing is due to needing to stabilize implementations. We believe the Llama 4 models are a significa…

There seems to be a lot of haloo, accusations, and rumors, but little meat to any of them. Maybe they rushed the release, were unsure of which one to go with, and some moderate rule bending in terms of which tune got sent to the arena, but I have seen no real hard evidence of real hard underhandedness.

Re: Meta got caught gaming AI benchmarks

#70

The truth is that the vast majority of FAANG engineers making high six figures are only good at deterministic work. They cant produce new things, and so meta and google are struggling to compete when actual merit matters, and they cant just brute force the solutions. Inside these companies, the massive tech systems built, are actually generally terrible, but they pile on legions of engineers to fix the problems. This…

> They cant produce new things, and so meta and google are struggling to compete when actual merit matters, and they cant just brute force the solutions.

You haven't been keeping up. Less than 2 weeks ago, Google released a model that has crushed the competition, clearly being SotA while currently effectively free for personal use.

Gemini 2.0 was already good, people just weren't paying attention. In fact 1.5 pro was already good, and ironically remains the #1 model at certain very specific tasks, despite being set for deprecation in September.

Google just suffered from their completely botched initial launch way back when (remember Bard?), rushed before the product was anywhere near ready, making them look lile a bunch of clowns compared to e.g. OpenAI. That left a lasting impression on those who don't devote significant time to keeping up with newer releases.

Post reply on HN