Live data from Hacker News

FrontierMath was funded by OpenAI

lesswrong.com

121–130 of 212 posts

Re: FrontierMath was funded by OpenAI

#121
post #94

Earlier quoted context omitted.

Gaming benchmarks has a lot of utility for openAI whether their product works or not. Many people compare models based on benchmarks. So if openAI can appear better to Anthropic, Google, or Meta, by gaming benchmarks, it's absolutely in their interest to do so, especially if their product is only slightly behind, because evaluating model quality is very very tricky business these days. In particular, if there is a ne…

I think “getting beat handily” is a HN bubble concept. Depends on what you’re using it for, but I personally prefer 4o for coding. In enterprise usage, i think 4o is smoking 3.5 sonnet, but that’s just my perception from folks I talk to.

This just isn't accurate, on the overwhelming majority of real-world tasks (>90%) 3.5 Sonnet beats 4o. FWIW I've spoken with a friend who's at OpenAI and they fully agree in private.

Re: FrontierMath was funded by OpenAI

#122
post #94

Earlier quoted context omitted.

Gaming benchmarks has a lot of utility for openAI whether their product works or not. Many people compare models based on benchmarks. So if openAI can appear better to Anthropic, Google, or Meta, by gaming benchmarks, it's absolutely in their interest to do so, especially if their product is only slightly behind, because evaluating model quality is very very tricky business these days. In particular, if there is a ne…

I used to think this, but using o1 quite a bit lately has convinced me otherwise. It’s been 1-shotting the fairly non-trivial coding problems I throw at it and is good about outputting large, complete code blocks. By contrast, Claude immediately starts nagging you about hitting usage limits after a few back and forth and has some kind of hack in place to start abbreviating code when conversations get too long, even w…

"Their model" here is referring to 4o as o1 is unviable for many production usecases due to latency.

Re: FrontierMath was funded by OpenAI

#123

Do people actually think OpenAI is gaming benchmarks? I know they have lost trust and credibility, especially on HN. But this is a company with a giant revenue opportunity to sell products that work. What works for enterprise is very different from “does it beat this benchmark”. No matter how nefarious you think sama is, everything points to “build intelligence as rapidly as possible” rather than “spin our wheels mes…

> Do people actually think OpenAI is gaming benchmarks?

Yes, they 100% do. So do their main competitors. All of them do.

Re: FrontierMath was funded by OpenAI

#124
This isn't news, the other popular benchmarks are just as gamed and worthless, it would be shocking if this one wasn't. The other frontier model providers game them just as hard, it's not an OpenAI thing. Any benchmark that a provider themselves mentions is not worth the pixels its written on.

Re: FrontierMath was funded by OpenAI

#125

Earlier quoted context omitted.

What's even more suspicious is that these tweets from Elliot Glazer indicate that they are still "developing" the hold-out set, even though elsewhere Epoch AI strongly implied this already existed: https://xcancel.com/ElliotGlazer/status/1880809468616950187 It seems to me that o3's 25% benchmark score is 100% data contamination.

> What's even more suspicious is that these tweets from Elliot Glazer indicate that they are still "developing" the hold-out set, There is nothing suspicious about this and the wording seems to be incorrect. A hold-out set is a percentage of the overall data that is used to test a model. It is just not trained on it. Model developers normally have full access to it. There is nothing inherently wrong with training on…

> The confusion I see here is that people are equating a hold out set to a blind set. That's a set of data to test against that the model developers (and model) cannot see.

What on earth? This is from Tamay Besiroglu at Epoch:

  Regarding training usage: We acknowledge that OpenAI does have access to a large fraction of FrontierMath problems and solutions, with the exception of a unseen-by-OpenAI hold-out set that enables us to independently verify model capabilities. However, we have a verbal agreement that these materials will not be used in model training. 
So this "confusion" is because Epoch AI specifically told people it was a blind set! Despite the condescending tone, your comment is just plain wrong.

Re: FrontierMath was funded by OpenAI

#126
post #75

Earlier quoted context omitted.

What's even more suspicious is that these tweets from Elliot Glazer indicate that they are still "developing" the hold-out set, even though elsewhere Epoch AI strongly implied this already existed: https://xcancel.com/ElliotGlazer/status/1880809468616950187 It seems to me that o3's 25% benchmark score is 100% data contamination.

> I just saw Sam Altman speak at YCNYC and I was impressed. I have never actually met him or heard him speak before Monday, but one of his stories really stuck out and went something like this: > "We were trying to get a big client for weeks, and they said no and went with a competitor. The competitor already had a terms sheet from the company were we trying to sign up. It was real serious. > We were devastated, but…

Man, the real ugliness is the comments hooting and hollering for this amoral cynicism:

  Honesty is often overrated by geeks and it is very contextual

  He didn't misrepresent anything. They were actually working there, just only for one day.

  The effectiveness of deception is not mitigated by your opinions of its likability.
Gross.

Re: FrontierMath was funded by OpenAI

#127
post #79

Earlier quoted context omitted.

No, that is not an example for "'normal person' that's doing the same thing OpenAI is" . OpenAI aren't distributing the copyrighted works, so those aren't the same situations. Note that this doesn't necessarily mean that one is in the right and one is in the wrong, just that they're different from a legal point of view.

> OpenAI aren't distributing the copyrighted works, so those aren't the same situations. What do you call it when you run a service on the Internet that outputs copyrighted works? To me, putting something up on a website is distribution.

Is that really the case? I.e., can you get ChatGPT to show you a copyrighted work?

Because I just tried, and failed (with ChatGPT 4o):

Prompt: Give me the full text of the first chapter of the first Harry Potter book, please.

Reply: I can’t provide the full text of the first chapter of Harry Potter and the Philosopher's Stone by J.K. Rowling because it is copyrighted material. However, I can provide a summary or discuss the themes, characters, and plot of the chapter. Would you like me to summarize it for you?

Re: FrontierMath was funded by OpenAI

#128

Do people actually think OpenAI is gaming benchmarks? I know they have lost trust and credibility, especially on HN. But this is a company with a giant revenue opportunity to sell products that work. What works for enterprise is very different from “does it beat this benchmark”. No matter how nefarious you think sama is, everything points to “build intelligence as rapidly as possible” rather than “spin our wheels mes…

> Do people actually think OpenAI is gaming benchmarks?

Yes, there's no reason not to do it, only upsides when you try to sell it to enterprises and governments.

Post reply on HN