Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

131–140 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#131
post #49

One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)

Why does AI have to be smarter than the collective of hummanity in order to be considered intelligent? It seems like we keep raising the bar on what intelligence means ¯\_(ツ)_/¯

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#132

Earlier quoted context omitted.

I've always assumed they removed it, because it's such a basic and fundamental part of ML training that you separate your test and train data. And yet I never see any papers even mention if/how they do this. And I wonder if they do, how do they guarantee with high reliability that their massive terabytes of data don't contain the answer.

First of all, Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly. Secondly, removing it from internet data is not 100% accurate. There are translations of the problems and solutions or references and direct match is not enough. MMLU and test set benchmarks show more resilience though in some previous research.

Let’s think about this.

> Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly

What exactly is the source of your belief that the Putnam would not be in the test data? Didn’t they train on everything they could get their hands on?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#133
post #49

One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)

Why does AI have to be smarter than the collective of hummanity in order to be considered intelligent? It seems like we keep raising the bar on what intelligence means ¯\_(ツ)_/¯

A machine that synthesizes all human knowledge really ought to know more than an individual in terms of intellect. An entity with all of human intellect prior to 1905 does not need to be as intelligent as a human to make discoveries that mere humans with limited intellect made. Why lower the bar?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#134

Earlier quoted context omitted.

First of all, Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly. Secondly, removing it from internet data is not 100% accurate. There are translations of the problems and solutions or references and direct match is not enough. MMLU and test set benchmarks show more resilience though in some previous research.

Let’s think about this. > Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly What exactly is the source of your belief that the Putnam would not be in the test data? Didn’t they train on everything they could get their hands on?

do you understand the difference between test data and train data? just reread this thread of comments

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#135

So many negative comments as if o3 didn’t get 25% on frontiermath - which is absolutely nuts. Sure, LLMs will perform better if the answer to a problem is directly in their training set. But that doesn’t mean they perform bad when the answer isn’t in their training set.

EpochAI have to send the questions (but not the answer key) to OpenAI in order to score the models. An overnight 2% -> 25% jump on this benchmark is a bit curious.

The 2% result belonged to a traditional LLM that costs cents to run, while o3 is extremely expensive.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#136
post #52

Or it's time to step back and call it what it is - very good pattern recognition. I mean, that's cool... we can get a lot of work done with pattern recognition. Most of the human race never really moves above that level of thinking in the workforce or navigating their daily life, especially if they default to various societally prescribed patterns of getting stuff done (eg. go to college or the military , find a job…

So, I am conflicted about this. If we take an example of what is considered a priori as creativity, such as story telling, LLMs can do pretty well at creating novel work. I can prompt with various parameters, plot elements, moral lessons, and get a de novo storyline, conflicts, relationships, character backstories, intrigues, and resolutions. Now, the writing style tends to be tone-deaf and poor at building tension f…

[deleted]

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#137
post #49

One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)

I had a similar thought but about asking the LLM to predict “future” major historical events. How much prompting would it take to predict wars, etc.?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#138

Earlier quoted context omitted.

I've always assumed they removed it, because it's such a basic and fundamental part of ML training that you separate your test and train data. And yet I never see any papers even mention if/how they do this. And I wonder if they do, how do they guarantee with high reliability that their massive terabytes of data don't contain the answer.

First of all, Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly. Secondly, removing it from internet data is not 100% accurate. There are translations of the problems and solutions or references and direct match is not enough. MMLU and test set benchmarks show more resilience though in some previous research.

funny that nobody replying to you seems to even know what a test set is. i always overestimate the depth of ML conversation you can have on HN

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#139
post #92

The paper includes several examples of their modified questions. There has been a substantial jump from o1-preview to o1, so I gave several samples to o1 and o1-pro ( not o1-preview), and current o1s gave the correct answer to those modified problems. SOTA changes fast.

The paper mentions that on several occasions the LLM will provide a correct answer but will either take big jumps without justifying them or will take illogical steps but end up with the right solution at the end. Did you check for that?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#140
post #49

One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)

> ask it for a formula for mass-energy equivalence

Way too easy. If you think that mass and energy might be equivalent, then dimensional analysis doesn’t give you too much choice in the formula. Really, the interesting thing about E=mc^2 isn’t the formula but the assertion that mass is a form of energy and all the surrounding observations about the universe.

Also, the actual insight in 1905 was more about asking the right questions and imagining that the equivalence principle could really hold, etc. A bunch of the math predates 1905 and would be there in an AI’s training set:

https://en.m.wikipedia.org/wiki/History_of_Lorentz_transform...

Post reply on HN