Live data from Hacker News

Meta got caught gaming AI benchmarks

theverge.com

161–170 of 171 posts

Re: Meta got caught gaming AI benchmarks

#161
post #54

Earlier quoted context omitted.

I don't know about Llama 4. Competition is intense in this field so you can't expect everybody to be number 1. However, I think the performance culture at Meta is counterproductive. Incentives are misaligned, I hope leadership will try to improve it. Employees are encouraged to ship half-baked features and move to another project. Quality isn't rewarded at all. The recent layoffs have made things even worse. Skilled…

I've never liked it, but > Move fast and break things is really a bad concept in this space, where you get limited shots at releasing something that generates interest. > Employees are encouraged to ship half-baked features And this is why I never liked that motto and have always pushed back at startups where I was hired that embraced this line of thought. Quality matters. It's context-dependent, so sometimes it matt…

I'm with you. Yet I've always understood 'move fast and break things' to mean that there is value in shipping stuff to production that is hard to obtain just sitting in the safe and relatively simple corner of your local development environment, polishing up things for eternity without any external feedback whatsoever, months or even years, and then doing a big drop. That's completely orthogonal to things like planning, testing, quality, taking time to think etc.

Maybe the fact that this is how I understood that motto is in itself telling.

Re: Meta got caught gaming AI benchmarks

#162
post #159

Earlier quoted context omitted.

Yes, their worst fear is people figuring out that an AI chatbot is a strict librarian that spits out quotes but doesn't let you enter the library (the AI model itself). Because with 3D game-like UIs people can enter the library and see all their stolen personal photos (if they were ever online), all kind of monsters. It'll be all over YouTube. Imagine this but you remove the noise and can walk like in an art gallery…

Do you really think facebook's model weights contain all the facebook personal photos?

With more than 50% probability. Instagram has a clause that they can use your photos for ads and endorsements, basically for profit.

Many sites repost those photos, I doubt Meta will care to meticulously remove them. If the model can generate photo-like content of people - it had photos of people put inside.

I doubt they only trained on public domain photos of people.

If they started doing shady stuff, why would they ever stop?

Re: Meta got caught gaming AI benchmarks

#165
post #23

Is LMArena junk now? I thought there was an aspect where you run two models on the same user-supplied query. Surely this can't be gamed? > “optimized for conversationality” I don't understand what that means - how it gives it an LMArena advantage.

It can be easily gamed. The users are self-selected, and they have zero incentive to be honest or rigorous or provide good responses. Some have incentives the opposite way. (There was a report of a prediction market user who said they had won a market on Gemini models by manipulating the votes; LMArena swore furiously there had definitely been no manipulation but was conspicuously silent on any details.) And the rele…

I guess I can't really refute your experience or observations. But, just a single anecdotal point; I use the arena's voting feature quite a bit and I try really hard to vote on the "best" answer. I've got no clue if the majority of the voters put the same level of effort into it or not, but I figure that an honest and rigorous vote is the least I can do in return for a service provided to me free with no obnoxious ads. There's a nonzero incentive to do the right thing, but it's hard to say where it comes from.

As an aside, I like getting two responses from two models that I can compare against one another (and with the primary sources of truth that I know a priori). Not only does that help me sanity-check the responses somewhat, but I get to interact with new models that I wouldn't have otherwise had the opportunity to. Learning new stuff is good, and being earnest is good.

_nick

Re: Meta got caught gaming AI benchmarks

#166
post #161
post #54

Earlier quoted context omitted.

I've never liked it, but > Move fast and break things is really a bad concept in this space, where you get limited shots at releasing something that generates interest. > Employees are encouraged to ship half-baked features And this is why I never liked that motto and have always pushed back at startups where I was hired that embraced this line of thought. Quality matters. It's context-dependent, so sometimes it matt…

I'm with you. Yet I've always understood 'move fast and break things' to mean that there is value in shipping stuff to production that is hard to obtain just sitting in the safe and relatively simple corner of your local development environment, polishing up things for eternity without any external feedback whatsoever, months or even years, and then doing a big drop. That's completely orthogonal to things like planni…

I understand it as that too. But even then, I dislike it.

Yes, "shipped, but terrible software" is far more valuable than "perfect software that's not available". But that's a goal.

The "move fast and break things" is a means to that goal. One of many ways to achieve this. In this, "move fast" is evident: who would ever want to move slowly if moving fast has the exact same trade-offs?

"Break things" is the part that I truly dislike. No. We don't "break things" for our users, our colleagues or our future-selves (ie tech debt).

In other situations, the "move fast and break things" implies a preferred leaning towards low-quality in the famous "Speed / Cost / Quality Trade-off". Where, by decreasing quality, we gain speed (while keeping cost the same?).

This fallacy has long been debunked, with actual research and data to back it up¹: Software Engineering projects that focus on quality, actually gain speed! Even without reading research or papers, this makes sense: we all know how much time is wasted on fixing regressions, muddling through technical debt, rewriting unreadable/-manageable/-extensible code etc. A clean, well-maintained, neatly tested, up-to-date codebase allows one to add features much faster than that horrible ball of mud that accumulated hacks and debt and bugs for decades.

¹e.g. Accelerate: The Science of Lean Software and DevOps: Building and Scaling High Performing Technology Organizations is a nice starting point with lots of references to research and data on this.

Re: Meta got caught gaming AI benchmarks

#167
post #23

Earlier quoted context omitted.

It can be easily gamed. The users are self-selected, and they have zero incentive to be honest or rigorous or provide good responses. Some have incentives the opposite way. (There was a report of a prediction market user who said they had won a market on Gemini models by manipulating the votes; LMArena swore furiously there had definitely been no manipulation but was conspicuously silent on any details.) And the rele…

I guess I can't really refute your experience or observations. But, just a single anecdotal point; I use the arena's voting feature quite a bit and I try really hard to vote on the "best" answer. I've got no clue if the majority of the voters put the same level of effort into it or not, but I figure that an honest and rigorous vote is the least I can do in return for a service provided to me free with no obnoxious ad…

In the time it takes you to do one good vote, the low-taste slopmaxxers can do 10 (or 20...?).

Re: Meta got caught gaming AI benchmarks

#168

Earlier quoted context omitted.

[flagged]

>They got the dataset from Epoch AI for one of the benchmarks and pinky swore that they wouldn't train on it Is anything here actually false or do you not like the conclusions that people may draw from it?

The false statement is “OpenAI got caught [gaming FrontierMath] a while back.”

From a primary source:

"OpenAI did not use FrontierMath data to guide the development of o1 or o3, at all.... we only downloaded FrontierMath for our evals long after the training data was frozen, and only looked at o3 FrontierMath results after the final announcement checkpoint was already picked."

https://x.com/__nmca__/status/1882563755806281986

Re: Meta got caught gaming AI benchmarks

#169
post #149

Earlier quoted context omitted.

[flagged]

[flagged]

The incorrect part is “OpenAI got caught [gaming FrontierMath] a while back.”

From a primary source:

"OpenAI did not use FrontierMath data to guide the development of o1 or o3, at all.... we only downloaded FrontierMath for our evals long after the training data was frozen, and only looked at o3 FrontierMath results after the final announcement checkpoint was already picked."

https://x.com/__nmca__/status/1882563755806281986

Re: Meta got caught gaming AI benchmarks

#170
post #140

Earlier quoted context omitted.

[flagged]

Why are you being disingenuous? Simply having access to the eval in question is already enough for your synthetics guys to match the distribution, and of course you don't contaminate directly on train, that would be stupid, and you would get caught, but if it does inform the reward, the result is the same. You _should_ quit, but you wouldn't because you'd already convinced yourself you're doing RL God's work, not sle…

From a primary source:

"OpenAI did not use FrontierMath data to guide the development of o1 or o3, at all.... we only downloaded FrontierMath for our evals long after the training data was frozen, and only looked at o3 FrontierMath results after the final announcement checkpoint was already picked."

https://x.com/__nmca__/status/1882563755806281986

Post reply on HN