Live data from Hacker News

FrontierMath was funded by OpenAI

lesswrong.com

131–140 of 212 posts

Re: FrontierMath was funded by OpenAI

#131

Earlier quoted context omitted.

> The "everybody does it" argument is a classic rationalization that doesn't actually justify anything. I'd argue here the more relevant point is "these specific people have been shown to have done it before." > The whole comment reads like someone who has picked up some ML terminology but lacks fundamental understanding of how research evaluation, technical accountability, and institutional incentives actually work…

> I'd argue here the more relevant point is "these specific people have been shown to have done it before." This is itself a slippery move. A vague gesture at past misconduct without actually specifying any incidents. If there's a clear pattern of documented benchmark manipulation, name it. Which benchmarks? When? What was the evidence? Without specifics, this is just trading one form of handwaving ("everyone does it…

This reply reads as though it were AI generated.

Let's bite though, and hope that unhelpful excessively long-winded replies are just your quirk.

> This is itself a slippery move. A vague gesture at past misconduct without actually specifying any incidents. If there's a clear pattern of documented benchmark manipulation, name it. Which benchmarks? When? What was the evidence? Without specifics, this is just trading one form of handwaving ("everyone does it") for another ("they did it before").

Ok, provide specifics yourself then. Someone replied and pointed out that they have every incentive to cheat, and your response was:

> This starts with a fallacious appeal to cynicism combined with an unsubstantiated claim about widespread misconduct. The "everybody does it" argument is a classic rationalization that doesn't actually justify anything. It also misunderstands the reputational and technical stakes - major labs face intense scrutiny of their methods and results, and there's plenty of incestuous movement between labs and plenty of leaks.

Respond to the content of the argument -- be specific. WHY is OpenAI not incentivized to cheat on this benchmark? Why is a once-nonprofit which turned from releasing open and transparent models to a closed model and begun raking in tens of billions of investor cash not incentivized to continue to make those investors happy? Be specific. Because there's a clear pattern of corporate behaviour at OpenAI and associated entities which suggests your take is not, in fact, the simpler viewpoint.

> This combines three questionable implications: > - That non-journal publications are automatically "not real concrete research" (tell that to physics/math arXiv)

Yes, arXiv will host lots of stuff that isn't real concrete research. They've hosted April Fool's jokes, for example.[1]

> - That peer review is binary - either traditional journal review or nothing (ignoring internal review processes, community peer review, public replications)

This is a poor/incorrect reading of the language. You have inferred meaning that does not exist. If citations are so important here, cite a few dozen that are peer reviewed out of the hundreds.

> - That volume ("pumped out") correlates with quality

Incorrect reading again. Volume here correlates with marketing and hype. It could have an effect on quality but that wasn't the purpose behind the language.

> You're making a valid critique of AI's departure from traditional academic structures, but then making an unjustified leap to assuming this means no rigor at all. It's like saying because a restaurant isn't Michelin-starred, it must have no food safety standards.

Why is that unjustified? It's no different than any of the science background people who have fallen into flat earther beliefs. They may understand the methods but if they are not tested with rigor and have abandoned scientific principles they do not get to keep pretending it's as valid as actual science.

> This also ignores the massive reputational and financial stakes that create strong incentives for internal rigor. Major labs have to maintain credibility with:

FWIW, this regurgitated talking point is what makes me believe this is an LLM-generated reply. OpenAI is not a major research lab. They appear to essentially to be trading off the names of more respected institutions and mathematicians who came up with FrontierMath. The credibility damage here can be done by a single person sharing data with OpenAI, unbeknownst to individual participants.

Separately, even under correct conditions it's not as if there are not all manner of problems in science in terms of ethical review. See for example, [2].

[1] https://arxiv.org/abs/2003.13879 - FWIW, I'm not against scientists having fun, but it should be understood that arXiv is basically three steps above HN or reddit. [2] https://lore.kernel.org/linux-nfs/YH+zwQgBBGUJdiVK@unreal/ + related HN discussion: https://news.ycombinator.com/item?id=26887670

Re: FrontierMath was funded by OpenAI

#132

Earlier quoted context omitted.

> an unsubstantiated claim about widespread misconduct. I can't prove it, but I heard it from multiple people in the industry. High contamination levels for existing benchmarks, though [1,2]. Whether to believe that it is just as good as we can do, not doing the best possible decontamination, or done on purpose is up to you. > Yes, validation and test sets serve different purposes - that's precisely why reputable lab…

> "I can't prove it, but I heard it from multiple people in the industry" The cited papers demonstrate that benchmark contamination exists as a general technical challenge, but are being misappropriated to support a much stronger claim about intentional misconduct by a specific actor. This is a textbook example of expanding evidence far, far, beyond its scope. > "The verbal agreement promised not to train on the eval…

> Attempting to justify potential misconduct through semantic technicalities ("well, validation isn't technically training")

Validation is not training, period. I'll ask again: what is the possible goal of accessing the evaluation set if you don't plan to use it for anything except the final evaluation, which is what the test set is used for? Either they just asked for access without any intent to use the provided data in any way except for final evaluation, which can be done without access, or they did somehow utilize the provided data, whether by training on it (which they verbally promised not to), using it as a validation set, using it to create a similar training set, or something else.

> This directly contradicts established principles of scientific integrity where the spirit of agreements matters as much as their letter.

OpenAI is not doing science; they are doing business.

> This represents a stark logical reversal. The initial argument assumed benchmark manipulation would be meaningful enough to influence investors and industry perception. Now, when challenged, the same metrics are suddenly "meaningless." This is fundamentally inconsistent - either the metrics matter (in which case manipulation would be serious misconduct) or they don't (in which case there's no incentive to manipulate them).

The metrics matter to people, but this doesn't mean people can meaningfully predict the model's performance using them. If I were trying to describe each of your arguments as some demagogue technique (you're going to call it ad hominem or something, probably), then I'd say it's a false dichotomy: it can, in fact, be impossible to use metrics to predict performance precisely enough and for people to care about metrics simultaneously.

> The attempted simultaneous appeal to and dismissal of credentials

I'm not appealing to credentials. Based on what I wrote, you made a wrong guess about my credentials, and I pointed out your mistake.

> at this point, the argument OpenAI did something rests on unfalsifiable claims about the industry as a whole, claiming insider knowledge, while avoiding any verifiable evidence.

Your position, on the other hand, rests on the assumption that corporations behave ethically and with integrity beyond what is required by the law (and, specifically, their contracts with other entities).

Re: FrontierMath was funded by OpenAI

#134

Earlier quoted context omitted.

OpenAI doesn't respect copyright so why would they let a verbal agreement get in the way of billion$

Can somehow explain to me how they can simply not respect copyright and get away with it? Also is this a uniquely open-ai problem, or also true of the other llm makers?

"When I was a kid, I was praying to a god for bicycle. But then I realized that god doesn't work this way, so I stole a bicycle and prayed to a god for forgiveness." (c)

Basically a heist too big and too fast to react. Now every impotent lawmaker in the world is afraid to call them what they are, because it will inflict on them wrath of both other IT corpos an of regular users, who will refuse to part with a toy they are now entitled to.

Re: FrontierMath was funded by OpenAI

#135

Earlier quoted context omitted.

>OpenAI has data access to much but not all of the dataset Their head mathematician says they have the full dataset, except a holdout set which they're currently developing (i.e. doesn't exist yet): https://www.reddit.com/r/singularity/comments/1i4n0r5/commen...

Thanks for the link. A holdout set which is yet to be used to verify the 25% claim. He also says that he doesn't believe that OpenAI would self-sabotage themselves by tricking the internal benchmarking performance since this will get easily exposed, either by the results from a holdout set or by the public repeating the benchmarks themselves. Seems reasonable to me.

>the public repeating the benchmarks themselves

The public has no access to this benchmark.

In fact, everyone thought it was all locked up in a vault at Epoch AI HQ, but looks like Sam Altman has a copy on his bedside table.

Re: FrontierMath was funded by OpenAI

#136

Earlier quoted context omitted.

Thanks for the link. A holdout set which is yet to be used to verify the 25% claim. He also says that he doesn't believe that OpenAI would self-sabotage themselves by tricking the internal benchmarking performance since this will get easily exposed, either by the results from a holdout set or by the public repeating the benchmarks themselves. Seems reasonable to me.

>the public repeating the benchmarks themselves The public has no access to this benchmark. In fact, everyone thought it was all locked up in a vault at Epoch AI HQ, but looks like Sam Altman has a copy on his bedside table.

Perhaps what he meant is that the public will be able to benchmark the model themselves by throwing different difficulty math problems at it and not necessarily the FrontierMath benchmark. It should become pretty obvious if they were faking the results or not.

Re: FrontierMath was funded by OpenAI

#137

Earlier quoted context omitted.

OpenAI doesn't respect copyright so why would they let a verbal agreement get in the way of billion$

Can somehow explain to me how they can simply not respect copyright and get away with it? Also is this a uniquely open-ai problem, or also true of the other llm makers?

I wonder if Google can sue them for downloading the YouTube videos plus automatically generated transcripts in order to train their models.

And if Google could enforce removal of this content from their training set and enforce a "rebuild" of a model which does not contain this data.

Billion-dollar lawsuits.

Re: FrontierMath was funded by OpenAI

#138
post #55

Earlier quoted context omitted.

Can somehow explain to me how they can simply not respect copyright and get away with it? Also is this a uniquely open-ai problem, or also true of the other llm makers?

Their argument is that using copyrighted data for training is transformative, and therefore a form of fair use. There are a number of ongoing lawsuits related to this issue, but so far the AI companies seem to be mostly winning. Eg. https://www.reuters.com/legal/litigation/openai-gets-partial... Some artists also tried to sue Stable Diffusion in Andersen v. Stability AI, and so far it looks like it's not going anywhe…

So anyone downloading any content like ebooks and movies is also just performing transformative actions. Forming memories, nothing else. Fair use.

Re: FrontierMath was funded by OpenAI

#139

Earlier quoted context omitted.

>the public repeating the benchmarks themselves The public has no access to this benchmark. In fact, everyone thought it was all locked up in a vault at Epoch AI HQ, but looks like Sam Altman has a copy on his bedside table.

Perhaps what he meant is that the public will be able to benchmark the model themselves by throwing different difficulty math problems at it and not necessarily the FrontierMath benchmark. It should become pretty obvious if they were faking the results or not.

It's been found [0] that slightly varying Putnam problems causes a 30% drop in o1-Preview accuracy, but that hasn't put a dent in OAI's hype.

There's absolutely no comeuppance for juicing benchmarks, especially ones no one has access to. If performance of o3 doesn't meet expectations, there'll be plenty of people making excuses for it ("You're prompting it wrong!", "That's just not its domain!").

[0] https://openreview.net/forum?id=YXnwlZe0yf&noteId=yrsGpHd0Sf

Re: FrontierMath was funded by OpenAI

#140

Earlier quoted context omitted.

Perhaps what he meant is that the public will be able to benchmark the model themselves by throwing different difficulty math problems at it and not necessarily the FrontierMath benchmark. It should become pretty obvious if they were faking the results or not.

It's been found [0] that slightly varying Putnam problems causes a 30% drop in o1-Preview accuracy, but that hasn't put a dent in OAI's hype. There's absolutely no comeuppance for juicing benchmarks, especially ones no one has access to. If performance of o3 doesn't meet expectations, there'll be plenty of people making excuses for it ("You're prompting it wrong!", "That's just not its domain!"). [0] https://openrevi…

> If performance of o3 doesn't meet expectations, there'll be plenty of people making excuses for it

I agree and I can definitely see that happening but it is also not impossible, given the incentive and impact of this technology, for some other company/community to create yet another, perhaps, FrontierMath-like benchmark to cross-validate the results.

I also don't disagree that it is not impossible for OpenAI to have faked these results. Time will tell.

Post reply on HN