Do people actually think OpenAI is gaming benchmarks? I know they have lost trust and credibility, especially on HN. But this is a company with a giant revenue opportunity to sell products that work. What works for enterprise is very different from “does it beat this benchmark”. No matter how nefarious you think sama is, everything points to “build intelligence as rapidly as possible” rather than “spin our wheels mes…
Gaming benchmarks has a lot of utility for openAI whether their product works or not. Many people compare models based on benchmarks. So if openAI can appear better to Anthropic, Google, or Meta, by gaming benchmarks, it's absolutely in their interest to do so, especially if their product is only slightly behind, because evaluating model quality is very very tricky business these days. In particular, if there is a ne…
FrontierMath was funded by OpenAI
101–110 of 212 posts
Re: FrontierMath was funded by OpenAI
#102What about model testing before releasing it?
Re: FrontierMath was funded by OpenAI
#103Do people actually think OpenAI is gaming benchmarks? I know they have lost trust and credibility, especially on HN. But this is a company with a giant revenue opportunity to sell products that work. What works for enterprise is very different from “does it beat this benchmark”. No matter how nefarious you think sama is, everything points to “build intelligence as rapidly as possible” rather than “spin our wheels mes…
Your fragrant disregard for ethics and focus on utilitarian aspects is certainly quite extreme to the extent that only a view people would agree with you in my view.
Re: FrontierMath was funded by OpenAI
#104My guess is that OpenAI didn't cheat as blatantly as just training on the test set. If they had, surely they could have gotten themselves an even higher mark than 25%. But I do buy the comment that they soft-cheated by using elements of the dataset for validation (which is absolutely still a form of data leakage). Even so, I suspect their reported number is roughly legit, because they report numbers on many benchmark…
> If they had, surely they could have gotten themselves an even higher mark than 25%. there is potentially some limitation of LLMs memorizing such complex proofs
Re: FrontierMath was funded by OpenAI
#105“… we have a verbal agreement that these materials will not be used in model training” Ha ha ha. Even written agreements are routinely violated as long as the potential upside > downside, and all you have is verbal agreement? And you didn’t disclose this? At the time o3 was released I wrote “this is so impressive that it brings out the pessimist in me”[0], thinking perhaps they were routing API calls to human workers…
Why would they use the materials in model training? It would defeat the purpose of having a benchmarking set
If you’re a for profit company trying to raise funding and fend off skepticism that your models really aren’t that much better than any one else’s, then…
It would be dishonest, but as long as no one found out until after you closed your funding round, there’s plenty of reason you might do this.
It comes down to caring about benchmarks and integrity or caring about piles of money.
Judge for yourself which one they chose.
Perhaps they didn’t train on it.
Who knows?
It’s fair to be skeptical though, under the circumstances.
Re: FrontierMath was funded by OpenAI
#106Why do people keep taking OpenAIs marketing spin at face value? This keeps happening, like when they neglected to mention that their most impressive Sora demo involved extensive manual editing/cleanup work because the studio couldn't get Sora to generate what they wanted. https://news.ycombinator.com/item?id=40359425
On each product they release, their top researchers are gradually leaving.
Everyone now knows what happens when you go against or question OpenAI after working for them, which is why you don't see any criticism and more of a cult-like worship.
Once again, "AGI" is a complete scam.
Re: FrontierMath was funded by OpenAI
#107Earlier quoted context omitted.
> the copyright holder’s tort would be with the source of the infringing distribution, not the people who read the material. Someone who just reads the material doesn't infringe. But someone who copies it, or prepares works that are derivative of it (which can happen even if they don't copy a single word or phrase literally), does. > would I then owe the authors of the books I learned from a fee to apply that knowled…
Right, but the onus of responsibility being on the end user publishing the song or creative work in violation of copyright, not the text editor, word processor, musical notation software, etc, correct? A text prediction tool isn’t a person, the data it is trained on is irrelevant to the copyright infringement perpetrated by the end user. They should perform due diligence to prevent liability.
Huh what? If a program "predicts" some data that is a derivative work of some copyrighted work (that the end user did not input), then ipso facto the tool itself is a derivative work of that copyrighted work, and illegal to distribute without permission. (Does that mean it's also illegal to publish and redistribute the brain of a human who's memorised a copyrighted work? Probably. I don't have a problem with that). How can it possibly be the user's responsibility when the user has never seen the copyrighted work being infringed on, only the software maker has?
And if you say that OpenAI isn't distributing their program but just offering it as a service, then we're back to the original situation: in that case OpenAI is illegally distributing derivative works of copyrighted works without permission. It's not even a YouTube like situation where some user uploaded the copyrighted work and they're just distributing it; OpenAI added the pirated books themselves.
Re: FrontierMath was funded by OpenAI
#108OpenAI played themselves here. Now nobody is going to take any of their results on this benchmark seriously, ever again. That o3 result has just disappeared in a poof of smoke. If they had blinded themselves properly then that wouldn't be the case. Whereas other AI companies now have the opportunity to be first to get a significant result on FrontierMath.
I refrain from forming a strong opinion in such situations. My intuition tells me that it's not cheating. But, well, it's intuition (probably based on my belief that the brain is nothing special physics-wise and it doesn't manage to realize unknown quantum algorithms in its warm and messy environment, so that classical computers can reproduce all of its feats when having appropriate algorithms and enough computing power. And math reasoning is just another step on a ladder of capabilities, not something that requires completely different approach). So, we'll see.
Re: FrontierMath was funded by OpenAI
#109Earlier quoted context omitted.
https://journa.host/@jeremiak/113811327999722586
No, that is not an example for "'normal person' that's doing the same thing OpenAI is" . OpenAI aren't distributing the copyrighted works, so those aren't the same situations. Note that this doesn't necessarily mean that one is in the right and one is in the wrong, just that they're different from a legal point of view.
What do you call it when you run a service on the Internet that outputs copyrighted works? To me, putting something up on a website is distribution.
Re: FrontierMath was funded by OpenAI
#110Earlier quoted context omitted.
The FSF funded some white papers a while ago on CoPilot: https://www.fsf.org/news/publication-of-the-fsf-funded-white... . Take a look at the analysis by two academics versed in law at https://www.fsf.org/licensing/copilot/copyright-implications... starting with §II.B that explains why it might be legal. Bradley Kuhn also has a differing opinion in another whitepaper there ( https://www.fsf.org/licensing/copilot/if-s…
All of the most capable models I use have been clearly trained on the entirety of libgen/z-lib. You know it is the first thing they did, it is like 100TB. Some of the models are even coy about it.