Live data from Hacker News

FrontierMath was funded by OpenAI

lesswrong.com

71–80 of 212 posts

Re: FrontierMath was funded by OpenAI

#71
post #27

My guess is that OpenAI didn't cheat as blatantly as just training on the test set. If they had, surely they could have gotten themselves an even higher mark than 25%. But I do buy the comment that they soft-cheated by using elements of the dataset for validation (which is absolutely still a form of data leakage). Even so, I suspect their reported number is roughly legit, because they report numbers on many benchmark…

> If they had, surely they could have gotten themselves an even higher mark than 25%.

there is potentially some limitation of LLMs memorizing such complex proofs

Re: FrontierMath was funded by OpenAI

#72
Even if OpenAI does not use these materials to directly train its models, OpenAI can collect or construct more data based on the knowledge points and test points of these questions to gain an unfair competitive advantage. It's like before the Gaokao, a teacher reads some of the Gaokao questions and then marks the test points in the book for you. This is cheating.

Re: FrontierMath was funded by OpenAI

#73
Many of these evals are quite easy to game. Often the actual evaluation part of benchmarking is left up to a good-faith actor, which was usually reasonable in academic settings less polluted by capital. AI labs, however, have disincentives to do a thorough or impartial job, so IMO we should never take their word for it. To verify, we need to be able to run these evals ourselves – this is only sometimes possible, as even if the datasets are public, the exact mechanisms of evaluation are not. In the long run, to be completely resilient to gaming via training, we probably need to follow suit of other fields and have third-party non-profit accredited (!!) evaluators who's entire premise is to evaluate, red-team, and generally keep AI safe and competent.

Re: FrontierMath was funded by OpenAI

#74

“… we have a verbal agreement that these materials will not be used in model training” Ha ha ha. Even written agreements are routinely violated as long as the potential upside > downside, and all you have is verbal agreement? And you didn’t disclose this? At the time o3 was released I wrote “this is so impressive that it brings out the pessimist in me”[0], thinking perhaps they were routing API calls to human workers…

OpenAI's benchmark results looking like Musk's Path of Exile character..

Re: FrontierMath was funded by OpenAI

#75

A co-founder of Epoch left a note in the comments: > We acknowledge that OpenAI does have access to a large fraction of FrontierMath problems and solutions, with the exception of a unseen-by-OpenAI hold-out set that enables us to independently verify model capabilities. However, we have a verbal agreement that these materials will not be used in model training. Ouch. A verbal agreement. As the saying goes, those aren…

What's even more suspicious is that these tweets from Elliot Glazer indicate that they are still "developing" the hold-out set, even though elsewhere Epoch AI strongly implied this already existed: https://xcancel.com/ElliotGlazer/status/1880809468616950187 It seems to me that o3's 25% benchmark score is 100% data contamination.

> I just saw Sam Altman speak at YCNYC and I was impressed. I have never actually met him or heard him speak before Monday, but one of his stories really stuck out and went something like this:

> "We were trying to get a big client for weeks, and they said no and went with a competitor. The competitor already had a terms sheet from the company were we trying to sign up. It was real serious.

> We were devastated, but we decided to fly down and sit in their lobby until they would meet with us. So they finally let us talk to them after most of the day.

> We then had a few more meetings, and the company wanted to come visit our offices so they could make sure we were a 'real' company. At that time, we were only 5 guys. So we hired a bunch of our college friends to 'work' for us for the day so we could look larger than we actually were. It worked, and we got the contract."

> I think the reason why PG respects Sam so much is he is charismatic, resourceful, and just overall seems like a genuine person.

https://news.ycombinator.com/item?id=3048944

Re: FrontierMath was funded by OpenAI

#76

Earlier quoted context omitted.

Can somehow explain to me how they can simply not respect copyright and get away with it? Also is this a uniquely open-ai problem, or also true of the other llm makers?

Simply put, if the model isn’t producing an actual copy, they aren’t violating copyright (in the US) under any current definition. As much as people bandy the term around, copyright has never applied to input, and the output of a tool is the responsibility of the end user. If I use a copy machine to reproduce your copyrighted work, I am responsible for that infringement not Xerox. If I coax your copyrighted work out…

> the copyright holder’s tort would be with the source of the infringing distribution, not the people who read the material.

Someone who just reads the material doesn't infringe. But someone who copies it, or prepares works that are derivative of it (which can happen even if they don't copy a single word or phrase literally), does.

> would I then owe the authors of the books I learned from a fee to apply that knowledge?

Facts can't be copyrighted, so applying the facts you learned is free, but creative works are generally copyrighted. If you write your own book inspired by a book you read, that can be copyright infringement (see The Wind Done Gone). If you use even a tiny fragment of someone else's work in your own, even if not consciously, that can be copyright infringement (see My Sweet Lord).

Re: FrontierMath was funded by OpenAI

#77

Earlier quoted context omitted.

OpenAI doesn't respect copyright so why would they let a verbal agreement get in the way of billion$

Can somehow explain to me how they can simply not respect copyright and get away with it? Also is this a uniquely open-ai problem, or also true of the other llm makers?

The short answer is that there is actually a number of active lawsuits alleging copyright violation, but they take time (years) to resolve. And since it's only been about two years since we've had the big generative AI blow up, fueled by entities with deep pockets (i.e., you can actually profit off of the lawsuit), there quite literally hasn't been enough time for a lawsuit to find them in violation of copyright.

And quite frankly, between the announcement of several licensing deals in the past year for new copyrighted content for training, and the recent decision in Warhol "clarifying" the definition of "transformative" for the purposes of fair use, the likelihood of training for AI being found fair is actually quite slim.

Re: FrontierMath was funded by OpenAI

#78

“… we have a verbal agreement that these materials will not be used in model training” Ha ha ha. Even written agreements are routinely violated as long as the potential upside > downside, and all you have is verbal agreement? And you didn’t disclose this? At the time o3 was released I wrote “this is so impressive that it brings out the pessimist in me”[0], thinking perhaps they were routing API calls to human workers…

Why would they use the materials in model training? It would defeat the purpose of having a benchmarking set

Re: FrontierMath was funded by OpenAI

#79
post #64

Earlier quoted context omitted.

Can you give an example of a copyright lawsuit lost by a 'normal person' that's doing the same thing OpenAI is?

https://journa.host/@jeremiak/113811327999722586

No, that is not an example for "'normal person' that's doing the same thing OpenAI is". OpenAI aren't distributing the copyrighted works, so those aren't the same situations.

Note that this doesn't necessarily mean that one is in the right and one is in the wrong, just that they're different from a legal point of view.

Re: FrontierMath was funded by OpenAI

#80
post #64

Earlier quoted context omitted.

Can you give an example of a copyright lawsuit lost by a 'normal person' that's doing the same thing OpenAI is?

https://journa.host/@jeremiak/113811327999722586

Aaron Swartz, while an infuriating tragedy, is antithetical to OpenAI's claim to transformation; he literally published documents that were behind a licensed paywall.
Post reply on HN