Earlier quoted context omitted.
Gaming benchmarks has a lot of utility for openAI whether their product works or not. Many people compare models based on benchmarks. So if openAI can appear better to Anthropic, Google, or Meta, by gaming benchmarks, it's absolutely in their interest to do so, especially if their product is only slightly behind, because evaluating model quality is very very tricky business these days. In particular, if there is a ne…
I think “getting beat handily” is a HN bubble concept. Depends on what you’re using it for, but I personally prefer 4o for coding. In enterprise usage, i think 4o is smoking 3.5 sonnet, but that’s just my perception from folks I talk to.
FrontierMath was funded by OpenAI
121–130 of 212 posts
Re: FrontierMath was funded by OpenAI
#122Earlier quoted context omitted.
Gaming benchmarks has a lot of utility for openAI whether their product works or not. Many people compare models based on benchmarks. So if openAI can appear better to Anthropic, Google, or Meta, by gaming benchmarks, it's absolutely in their interest to do so, especially if their product is only slightly behind, because evaluating model quality is very very tricky business these days. In particular, if there is a ne…
I used to think this, but using o1 quite a bit lately has convinced me otherwise. It’s been 1-shotting the fairly non-trivial coding problems I throw at it and is good about outputting large, complete code blocks. By contrast, Claude immediately starts nagging you about hitting usage limits after a few back and forth and has some kind of hack in place to start abbreviating code when conversations get too long, even w…
Re: FrontierMath was funded by OpenAI
#123Do people actually think OpenAI is gaming benchmarks? I know they have lost trust and credibility, especially on HN. But this is a company with a giant revenue opportunity to sell products that work. What works for enterprise is very different from “does it beat this benchmark”. No matter how nefarious you think sama is, everything points to “build intelligence as rapidly as possible” rather than “spin our wheels mes…
Yes, they 100% do. So do their main competitors. All of them do.
Re: FrontierMath was funded by OpenAI
#124Re: FrontierMath was funded by OpenAI
#125Earlier quoted context omitted.
What's even more suspicious is that these tweets from Elliot Glazer indicate that they are still "developing" the hold-out set, even though elsewhere Epoch AI strongly implied this already existed: https://xcancel.com/ElliotGlazer/status/1880809468616950187 It seems to me that o3's 25% benchmark score is 100% data contamination.
> What's even more suspicious is that these tweets from Elliot Glazer indicate that they are still "developing" the hold-out set, There is nothing suspicious about this and the wording seems to be incorrect. A hold-out set is a percentage of the overall data that is used to test a model. It is just not trained on it. Model developers normally have full access to it. There is nothing inherently wrong with training on…
What on earth? This is from Tamay Besiroglu at Epoch:
Regarding training usage: We acknowledge that OpenAI does have access to a large fraction of FrontierMath problems and solutions, with the exception of a unseen-by-OpenAI hold-out set that enables us to independently verify model capabilities. However, we have a verbal agreement that these materials will not be used in model training.
So this "confusion" is because Epoch AI specifically told people it was a blind set! Despite the condescending tone, your comment is just plain wrong.Re: FrontierMath was funded by OpenAI
#126Earlier quoted context omitted.
What's even more suspicious is that these tweets from Elliot Glazer indicate that they are still "developing" the hold-out set, even though elsewhere Epoch AI strongly implied this already existed: https://xcancel.com/ElliotGlazer/status/1880809468616950187 It seems to me that o3's 25% benchmark score is 100% data contamination.
> I just saw Sam Altman speak at YCNYC and I was impressed. I have never actually met him or heard him speak before Monday, but one of his stories really stuck out and went something like this: > "We were trying to get a big client for weeks, and they said no and went with a competitor. The competitor already had a terms sheet from the company were we trying to sign up. It was real serious. > We were devastated, but…
Honesty is often overrated by geeks and it is very contextual
He didn't misrepresent anything. They were actually working there, just only for one day.
The effectiveness of deception is not mitigated by your opinions of its likability.
Gross.Re: FrontierMath was funded by OpenAI
#127Earlier quoted context omitted.
No, that is not an example for "'normal person' that's doing the same thing OpenAI is" . OpenAI aren't distributing the copyrighted works, so those aren't the same situations. Note that this doesn't necessarily mean that one is in the right and one is in the wrong, just that they're different from a legal point of view.
> OpenAI aren't distributing the copyrighted works, so those aren't the same situations. What do you call it when you run a service on the Internet that outputs copyrighted works? To me, putting something up on a website is distribution.
Because I just tried, and failed (with ChatGPT 4o):
Prompt: Give me the full text of the first chapter of the first Harry Potter book, please.
Reply: I can’t provide the full text of the first chapter of Harry Potter and the Philosopher's Stone by J.K. Rowling because it is copyrighted material. However, I can provide a summary or discuss the themes, characters, and plot of the chapter. Would you like me to summarize it for you?
Re: FrontierMath was funded by OpenAI
#128Do people actually think OpenAI is gaming benchmarks? I know they have lost trust and credibility, especially on HN. But this is a company with a giant revenue opportunity to sell products that work. What works for enterprise is very different from “does it beat this benchmark”. No matter how nefarious you think sama is, everything points to “build intelligence as rapidly as possible” rather than “spin our wheels mes…
Yes, there's no reason not to do it, only upsides when you try to sell it to enterprises and governments.
Re: FrontierMath was funded by OpenAI
#129Why does it have a customer service popover chat assistant?
Re: FrontierMath was funded by OpenAI
#130Man, this is huge.