Live data from Hacker News

FrontierMath was funded by OpenAI

lesswrong.com

111–120 of 212 posts

Re: FrontierMath was funded by OpenAI

#111

Earlier quoted context omitted.

Why wouldn't OpenAI cheat? It's an open secret in industry that benchmarks are trained on. Everybody does it, so you need to do that or else your similarly performing model will look worse on paper. And even they respect the agreement, even using test set as a validation set can be a huge advantage. That's why validation set and test set are two different terms with precise meaning. As for "knowing it's bad", most pe…

> Why wouldn't OpenAI cheat? It's an open secret in industry that benchmarks are trained on. Everybody does it, so you need to do that or else your similarly performing model will look worse on paper. This starts with a fallacious appeal to cynicism combined with an unsubstantiated claim about widespread misconduct. The "everybody does it" argument is a classic rationalization that doesn't actually justify anything.…

> The "everybody does it" argument is a classic rationalization that doesn't actually justify anything.

I'd argue here the more relevant point is "these specific people have been shown to have done it before."

> The whole comment reads like someone who has picked up some ML terminology but lacks fundamental understanding of how research evaluation, technical accountability, and institutional incentives actually work in the field. The dismissive tone and casual accusations of misconduct don't help their credibility either.

I think what you're missing is the observation that so very little of that is actually applied in this case. "AI" here is not being treated as an actual science would be. The majority of the papers pumped out of these places are not real concrete research, not submitted to journals, and not peer reviewed works.

Re: FrontierMath was funded by OpenAI

#112
post #6

> Tamay from Epoch AI here. We made a mistake in not being more transparent about OpenAI's involvement. We were restricted from disclosing the partnership until around the time o3 launched, and in hindsight we should have negotiated harder for the ability to be transparent to the benchmark contributors as soon as possible. Our contract specifically prevented us from disclosing information about the funding source and…

>OpenAI has data access to much but not all of the dataset Their head mathematician says they have the full dataset, except a holdout set which they're currently developing (i.e. doesn't exist yet): https://www.reddit.com/r/singularity/comments/1i4n0r5/commen...

Thanks for the link. A holdout set which is yet to be used to verify the 25% claim. He also says that he doesn't believe that OpenAI would self-sabotage themselves by tricking the internal benchmarking performance since this will get easily exposed, either by the results from a holdout set or by the public repeating the benchmarks themselves. Seems reasonable to me.

Re: FrontierMath was funded by OpenAI

#113

Earlier quoted context omitted.

> Why wouldn't OpenAI cheat? It's an open secret in industry that benchmarks are trained on. Everybody does it, so you need to do that or else your similarly performing model will look worse on paper. This starts with a fallacious appeal to cynicism combined with an unsubstantiated claim about widespread misconduct. The "everybody does it" argument is a classic rationalization that doesn't actually justify anything.…

> The "everybody does it" argument is a classic rationalization that doesn't actually justify anything. I'd argue here the more relevant point is "these specific people have been shown to have done it before." > The whole comment reads like someone who has picked up some ML terminology but lacks fundamental understanding of how research evaluation, technical accountability, and institutional incentives actually work…

> I'd argue here the more relevant point is "these specific people have been shown to have done it before."

This is itself a slippery move. A vague gesture at past misconduct without actually specifying any incidents. If there's a clear pattern of documented benchmark manipulation, name it. Which benchmarks? When? What was the evidence? Without specifics, this is just trading one form of handwaving ("everyone does it") for another ("they did it before").

> "AI" here is not being treated as an actual science would be.

There's some truth here but also some sleight of hand. Yes, AI development often moves outside traditional academic channels. But, you imply this automatically means less rigor, which doesn't follow. Many industry labs have internal review processes, replication requirements, and validation procedures that can be as or more stringent than academic peer review. The fact that something isn't in Nature doesn't automatically make it less rigorous.

> The majority of the papers pumped out of these places are not real concrete research, not submitted to journals, and not peer reviewed works.

This combines three questionable implications:

- That non-journal publications are automatically "not real concrete research" (tell that to physics/math arXiv)

- That peer review is binary - either traditional journal review or nothing (ignoring internal review processes, community peer review, public replications)

- That volume ("pumped out") correlates with quality

You're making a valid critique of AI's departure from traditional academic structures, but then making an unjustified leap to assuming this means no rigor at all. It's like saying because a restaurant isn't Michelin-starred, it must have no food safety standards.

This also ignores the massive reputational and financial stakes that create strong incentives for internal rigor. Major labs have to maintain credibility with:

- Their own employees.

- Other researchers who will try to replicate results.

- Partners integrating their technology.

- Investors doing technical due diligence.

- Regulators scrutinizing their claims.

The idea that they would casually risk all that just to bump up one benchmark number (but not too much! just from 10% to 35%) doesn't align with the actual incentive structure these organizations face.

Both the original comment and this fall into the same trap - mistaking cynicism for sophistication while actually displaying a somewhat superficial understanding of how modern AI research and development actually operates.

Re: FrontierMath was funded by OpenAI

#114

Earlier quoted context omitted.

Why wouldn't OpenAI cheat? It's an open secret in industry that benchmarks are trained on. Everybody does it, so you need to do that or else your similarly performing model will look worse on paper. And even they respect the agreement, even using test set as a validation set can be a huge advantage. That's why validation set and test set are two different terms with precise meaning. As for "knowing it's bad", most pe…

> Why wouldn't OpenAI cheat? It's an open secret in industry that benchmarks are trained on. Everybody does it, so you need to do that or else your similarly performing model will look worse on paper. This starts with a fallacious appeal to cynicism combined with an unsubstantiated claim about widespread misconduct. The "everybody does it" argument is a classic rationalization that doesn't actually justify anything.…

> an unsubstantiated claim about widespread misconduct.

I can't prove it, but I heard it from multiple people in the industry. High contamination levels for existing benchmarks, though [1,2]. Whether to believe that it is just as good as we can do, not doing the best possible decontamination, or done on purpose is up to you.

> Yes, validation and test sets serve different purposes - that's precisely why reputable labs maintain strict separations between them.

The verbal agreement promised not to train on the evaluation set. Using it as a validation set would not violate this agreement. Clearly, OpenAI did not plan to use the provided evaluation as a testset, because then they wouldn't need access to it. Also, reporting validation numbers as performance metric is not unheard of.

> This reveals a fundamental misunderstanding of why math capabilities matter. They're not primarily about serving math users - they're a key metric for abstract reasoning and systematic problem-solving abilities.

How good of a proxy is it? There is some correlation, but can you say something quantitative? Do you think you can predict which models perform better on math benchmarks based on interaction with them? Especially for a benchmark you have no access to and can't solve by yourself? If the answer is no, the number is more or less meaningless by itself, which means it would be very hard to catch somebody giving you incorrect numbers.

> someone who has picked up some ML terminology but lacks fundamental understanding of how research evaluation, technical accountability, and institutional incentives actually work in the field

My credentials are in my profile, not that I think they should matter. However, I do have experience specifically in deep learning research and evaluation of LLMs.

[1] https://aclanthology.org/2024.naacl-long.482/ [2] https://arxiv.org/abs/2412.15194

Re: FrontierMath was funded by OpenAI

#115

“… we have a verbal agreement that these materials will not be used in model training” Ha ha ha. Even written agreements are routinely violated as long as the potential upside > downside, and all you have is verbal agreement? And you didn’t disclose this? At the time o3 was released I wrote “this is so impressive that it brings out the pessimist in me”[0], thinking perhaps they were routing API calls to human workers…

Not used in model training probably means it was used in model validation.

Re: FrontierMath was funded by OpenAI

#116
“we now know how to build AGI” --Sam Altman.

which should really be “we now know how to improve associative reasoning but we still need to cheat when it comes to math because the bottom line is that the models can only capture logic associatively, not synthesize deductively, which is what’s needed for math beyond recipe-based reasoning"

Re: FrontierMath was funded by OpenAI

#117

Earlier quoted context omitted.

> Why wouldn't OpenAI cheat? It's an open secret in industry that benchmarks are trained on. Everybody does it, so you need to do that or else your similarly performing model will look worse on paper. This starts with a fallacious appeal to cynicism combined with an unsubstantiated claim about widespread misconduct. The "everybody does it" argument is a classic rationalization that doesn't actually justify anything.…

> an unsubstantiated claim about widespread misconduct. I can't prove it, but I heard it from multiple people in the industry. High contamination levels for existing benchmarks, though [1,2]. Whether to believe that it is just as good as we can do, not doing the best possible decontamination, or done on purpose is up to you. > Yes, validation and test sets serve different purposes - that's precisely why reputable lab…

> "I can't prove it, but I heard it from multiple people in the industry"

The cited papers demonstrate that benchmark contamination exists as a general technical challenge, but are being misappropriated to support a much stronger claim about intentional misconduct by a specific actor. This is a textbook example of expanding evidence far, far, beyond its scope.

> "The verbal agreement promised not to train on the evaluation set. Using it as a validation set would not violate this agreement."

This argument reveals a concerning misunderstanding of research ethics. Attempting to justify potential misconduct through semantic technicalities ("well, validation isn't technically training") suggests a framework where anything not explicitly forbidden is acceptable. This directly contradicts established principles of scientific integrity where the spirit of agreements matters as much as their letter.

> "How good of a proxy is it? [...] If the answer is no, the number is more or less meaningless by itself"

This represents a stark logical reversal. The initial argument assumed benchmark manipulation would be meaningful enough to influence investors and industry perception. Now, when challenged, the same metrics are suddenly "meaningless." This is fundamentally inconsistent - either the metrics matter (in which case manipulation would be serious misconduct) or they don't (in which case there's no incentive to manipulate them).

> "My credentials are in my profile, not that I think they should matter."

The attempted simultaneous appeal to and dismissal of credentials is an interesting mirror of the claims as a whole: at this point, the argument OpenAI did something rests on unfalsifiable claims about the industry as a whole, claiming insider knowledge, while avoiding any verifiable evidence.

When challenged, it retreats to increasingly abstract hypotheticals about what "could" happen rather than what evidence shows did happen.

This demonstrates how seemingly technical arguments can fail basic principles of evidence and logic, while maintaining surface-level plausibility through domain-specific terminology. This kind of reasoning would not pass basic scrutiny in any rigorous research context.

Re: FrontierMath was funded by OpenAI

#118
post #7

if they used it in training it should be 100% hit. most likely they used it to verify and tune parameters.

> if they used it in training it should be 100% hit. Not necessarily, no. A statistical model will attempt to minimise overall loss, generally speaking. If it gets 100% accuracy on the training data it's usually an overfit. (Hugging the data points too tightly, thereby failing to predict real life cases)

you are mostly right. but seeing almost perfectly reconstructed images from training set it's obvious model -can- memorize samples. in this case it would reproduce the answers too close to the original to be just 'accidental'. should be easy to test.

My guess samples could be used to find good enough stopping point for o1, o3 models. which is hardcoded.

Re: FrontierMath was funded by OpenAI

#119
So in conclusion, any evaluation of openai models on frontier math is increadibly invalidated.

I would even go so far as to say this invalidates not only FrontierMath but also anything Epoch AI has and will touch.

Any academic misjudgement like this massive conflict and cheating makes you unthrustworthy in a academic context.

Re: FrontierMath was funded by OpenAI

#120

Do people actually think OpenAI is gaming benchmarks? I know they have lost trust and credibility, especially on HN. But this is a company with a giant revenue opportunity to sell products that work. What works for enterprise is very different from “does it beat this benchmark”. No matter how nefarious you think sama is, everything points to “build intelligence as rapidly as possible” rather than “spin our wheels mes…

Yes, it looks all but certain that OpenAI gamed this particular benchmark.

Otherwise, they would not have had a contract that prohibited revealing that OpenAI was involved with the project until after the o3 announcements were made and the market had time to react. There is no reason to have such a specific agreement unless you plan to use the backdoor access to beat the benchmark: otherwise, OpenAI would not have known in advance that o3 will perform well! In fact, if there was proper blinding in place (which Epoch heads confirmed was not the case), there would have been no reason for secrecy at all.

Google, xAI and Anthropic's test-time compute experiments were really underwhelming: if OpenAI has secret access to benchmarks, that explains why their performance is so different.

Post reply on HN