Live data from Hacker News

Meta got caught gaming AI benchmarks

theverge.com

71–80 of 171 posts

Re: Meta got caught gaming AI benchmarks

#71

The truth is that the vast majority of FAANG engineers making high six figures are only good at deterministic work. They cant produce new things, and so meta and google are struggling to compete when actual merit matters, and they cant just brute force the solutions. Inside these companies, the massive tech systems built, are actually generally terrible, but they pile on legions of engineers to fix the problems. This…

A whole lot of opinion there, not a whole lot of evidence.

Re: Meta got caught gaming AI benchmarks

#72
post #70

The truth is that the vast majority of FAANG engineers making high six figures are only good at deterministic work. They cant produce new things, and so meta and google are struggling to compete when actual merit matters, and they cant just brute force the solutions. Inside these companies, the massive tech systems built, are actually generally terrible, but they pile on legions of engineers to fix the problems. This…

> They cant produce new things, and so meta and google are struggling to compete when actual merit matters, and they cant just brute force the solutions. You haven't been keeping up. Less than 2 weeks ago, Google released a model that has crushed the competition, clearly being SotA while currently effectively free for personal use. Gemini 2.0 was already good, people just weren't paying attention. In fact 1.5 pro was…

gemini 2.5 pro isnt good, and if you think it is, you arent using LLMs correctly. The model gets crushed by o1 pro and sonnet 3.7 thinking. Build a large contextual prompt ( > 50k tokens) with a ton of code, and see how bad it is. I cancelled my gemini subscription

Re: Meta got caught gaming AI benchmarks

#73

The truth is that the vast majority of FAANG engineers making high six figures are only good at deterministic work. They cant produce new things, and so meta and google are struggling to compete when actual merit matters, and they cant just brute force the solutions. Inside these companies, the massive tech systems built, are actually generally terrible, but they pile on legions of engineers to fix the problems. This…

The problem is less that those high level engineers are only good at deterministic work and more that they're only rewarded for deterministic work.

There is no system to pitch an idea as opening new frontiers - all ideas must be able to optimize some number that leadership has already been tricked into believing is important.

Re: Meta got caught gaming AI benchmarks

#74
post #34

I think it's most illustrative to see the sample battles (H2H) that LMArena released [1]. The outputs of Meta's model is too verbose and too 'yappy' IMO. And looking at the verdicts, it's no wonder by people are discounting LMArena rankings. [1]: https://huggingface.co/spaces/lmarena-ai/Llama-4-Maverick-03...

Yep, it’s clear that many wins are due to Llama 4’s lowered refusal rate which is an effective form of elo hacking.

Re: Meta got caught gaming AI benchmarks

#75

The truth is that the vast majority of FAANG engineers making high six figures are only good at deterministic work. They cant produce new things, and so meta and google are struggling to compete when actual merit matters, and they cant just brute force the solutions. Inside these companies, the massive tech systems built, are actually generally terrible, but they pile on legions of engineers to fix the problems. This…

A whole lot of opinion there, not a whole lot of evidence.

Evidence is a decade inside these companies, watching the circus

Re: Meta got caught gaming AI benchmarks

#76

I tried to make Studio Ghibli inspired images using presumably their new models. It was ass.

GPT 4o images is the future of all image gen.

Every other player: Black Forest Labs' Flux, Stability.ai's Stable Diffusion, and even closed models like Ideogram and Midjourney, are all on the path to extinction.

Image generation and editing must be multimodal. Full stop.

Google Imagen will probably be the first model to match the capabilities of 4o. I'm hoping one of the open weights labs or Chinese AI giants will release a model that demonstrates similar capabilities soon. That'll keep the race neck and neck.

Re: Meta got caught gaming AI benchmarks

#77

The truth is that the vast majority of FAANG engineers making high six figures are only good at deterministic work. They cant produce new things, and so meta and google are struggling to compete when actual merit matters, and they cant just brute force the solutions. Inside these companies, the massive tech systems built, are actually generally terrible, but they pile on legions of engineers to fix the problems. This…

[deleted]

Re: Meta got caught gaming AI benchmarks

#79

Earlier quoted context omitted.

I don't know about Llama 4. Competition is intense in this field so you can't expect everybody to be number 1. However, I think the performance culture at Meta is counterproductive. Incentives are misaligned, I hope leadership will try to improve it. Employees are encouraged to ship half-baked features and move to another project. Quality isn't rewarded at all. The recent layoffs have made things even worse. Skilled…

I agree. I think of it like a car engine. You can push it up to a certain RPM and it will keep making more and more power. Above that RPM, the engine starts to produce less power and eventually blows a gasket. I think the performance-based management worked for a while because there were some gains to be had by pushing people harder. However, they’ve gone past that and are now pushing people too hard and getting wors…

[deleted]
Post reply on HN