Live data from Hacker News

Meta got caught gaming AI benchmarks

theverge.com

91–100 of 171 posts

Re: Meta got caught gaming AI benchmarks

#91
post #76

Earlier quoted context omitted.

GPT 4o images is the future of all image gen. Every other player: Black Forest Labs' Flux, Stability.ai's Stable Diffusion, and even closed models like Ideogram and Midjourney, are all on the path to extinction. Image generation and editing must be multimodal. Full stop. Google Imagen will probably be the first model to match the capabilities of 4o. I'm hoping one of the open weights labs or Chinese AI giants will re…

One very important distinction between image models is the implementation: 4o is autogressive, slow, and extremely expensive. Although the Ghibli trend is market validation, I suspect that competitors may not want to copy it just yet.

> 4o is autogressive, slow, and extremely expensive.

If you factor in the amount of time wasted with prompting and inpainting, it's extremely well worth it.

Re: Meta got caught gaming AI benchmarks

#92
post #70

Earlier quoted context omitted.

> They cant produce new things, and so meta and google are struggling to compete when actual merit matters, and they cant just brute force the solutions. You haven't been keeping up. Less than 2 weeks ago, Google released a model that has crushed the competition, clearly being SotA while currently effectively free for personal use. Gemini 2.0 was already good, people just weren't paying attention. In fact 1.5 pro was…

gemini 2.5 pro isnt good, and if you think it is, you arent using LLMs correctly. The model gets crushed by o1 pro and sonnet 3.7 thinking. Build a large contextual prompt ( > 50k tokens) with a ton of code, and see how bad it is. I cancelled my gemini subscription

I have, dozens of times, and it's generally better than 3.7. Especially with more context it's less forgetful. o1-pro is absurdly expensive and slow, good luck using that with tools. Virtually all benchmarks, including less gamed ones such as Aider's, show the same. WebLM still has 3.7 ahead, with Sonnet always having been particularly strong at web development, but even on there 2.5 Pro is miles in front of any OpenAI model.

Gemini subscription? Surely if you're "using LLMs correctly" you'd have been using the APIs for everything anyway. Subscriptions are generally for non-techy consumers.

In any case, just straight up saying "it isn't good" is absurd, even if you personally prefer others.

Re: Meta got caught gaming AI benchmarks

#93
post #22

Earlier quoted context omitted.

Sarcasm doesn't translate well in text. Please, elaborate.

Matt Levine has a common refrain which is that basically everything is securities fraud. If you were an investor who invested on the premise that Meta was good at AI and Zuck knowingly put out a bad model, is that securities fraud? Matt Levine will probably argue that it could be in a future edition of Money Stuff (his very good newsletter).

The "everything is securities fraud" meme is really unfortunate, not quite as bad as the "fiduciary duty means execs have to chase short-term profit" myth, but still harmful.

It's only because lying ("puffery") about everything has become the norm in corporate America that indeed, almost all listed companies commit securities fraud. If they'd go back to being honest businessmen, no more securities fraud. Just stop claiming things that aren't true. This is a very real option they could take. If they don't, then they're willingly and knowingly commiting securities fraud. But the meme makes it sound to people as if it's unavoidable, when it's anything but.

Re: Meta got caught gaming AI benchmarks

#94
post #6

Earlier quoted context omitted.

Not even first, OpenAI got caught a while back

People have been gaming ML benchmarks as long as there have been ML benchmarks. That's why it's better to see if other researchers are incorporating a technique into their actual models rather than 'is this paper the bold entry in a benchmark table'. But it takes longer.

When a measure becomes a target it is no longer a good measure.

These ML benchmarks were never going to last very long. There is far too much pressure to game them, even unintentionally.

Re: Meta got caught gaming AI benchmarks

#95
Impressive results from Meta's Llama adapting to various benchmarks. However, gaming performance seems lackluster compared to specialized models like Alpaca. It raises questions about the viability of large language models for complex, interactive tasks like gaming without more targeted fine-tuning. Exciting progress nonetheless!

Re: Meta got caught gaming AI benchmarks

#96
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

"made an ambitious bet on MoEs"? No, DeepSeek is MoE, and they succeeded. Meta is not betting on MoE, it just does what other people have done.

Llama4 seems in many ways a cut and paste of DeepSeek. Including the shared expert and the high sparsity. It's a DeepSeek that does not work well.

Re: Meta got caught gaming AI benchmarks

#97

I tried to make Studio Ghibli inspired images using presumably their new models. It was ass.

Llama is not an image generating model. Any interface that uses Llama and generates images is calling out to a separate image generator as a tool, like OpenAI used to do with ChatGPT and DALL-E up until a couple of weeks ago: https://simonwillison.net/2023/Oct/26/add-a-walrus/

Re: Meta got caught gaming AI benchmarks

#98
post #9

The Llama 4 launch looks like a real debacle for Meta. The model doesn't look great. All the coverage I've seen has been negative. This is about what I expected, but it makes you wonder what they're going to do next. At this point it looks like they are falling behind the other open models, and made an ambitious bet on MoEs, without this paying off. Did Zuck push for the release? I'm sure they knew it wasn't ready ye…

[dead]

Re: Meta got caught gaming AI benchmarks

#99
post #76

I tried to make Studio Ghibli inspired images using presumably their new models. It was ass.

GPT 4o images is the future of all image gen. Every other player: Black Forest Labs' Flux, Stability.ai's Stable Diffusion, and even closed models like Ideogram and Midjourney, are all on the path to extinction. Image generation and editing must be multimodal. Full stop. Google Imagen will probably be the first model to match the capabilities of 4o. I'm hoping one of the open weights labs or Chinese AI giants will re…

[deleted]

Re: Meta got caught gaming AI benchmarks

#100
post #12

tech companies competing over something that is losing them money is the most bizarre spectacle yet.

The reason is simple. All tech companies have very high valuations. They have to sell investors a dream to justify that valuation. They have to convince people that they have the next big thing around the corner.
Post reply on HN