One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)
30% drop in O1-preview accuracy when Putnam problems are slightly variated
131–140 of 558 posts
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#132Earlier quoted context omitted.
I've always assumed they removed it, because it's such a basic and fundamental part of ML training that you separate your test and train data. And yet I never see any papers even mention if/how they do this. And I wonder if they do, how do they guarantee with high reliability that their massive terabytes of data don't contain the answer.
First of all, Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly. Secondly, removing it from internet data is not 100% accurate. There are translations of the problems and solutions or references and direct match is not enough. MMLU and test set benchmarks show more resilience though in some previous research.
> Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly
What exactly is the source of your belief that the Putnam would not be in the test data? Didn’t they train on everything they could get their hands on?
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#133One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)
Why does AI have to be smarter than the collective of hummanity in order to be considered intelligent? It seems like we keep raising the bar on what intelligence means ¯\_(ツ)_/¯
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#134Earlier quoted context omitted.
First of all, Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly. Secondly, removing it from internet data is not 100% accurate. There are translations of the problems and solutions or references and direct match is not enough. MMLU and test set benchmarks show more resilience though in some previous research.
Let’s think about this. > Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly What exactly is the source of your belief that the Putnam would not be in the test data? Didn’t they train on everything they could get their hands on?
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#135So many negative comments as if o3 didn’t get 25% on frontiermath - which is absolutely nuts. Sure, LLMs will perform better if the answer to a problem is directly in their training set. But that doesn’t mean they perform bad when the answer isn’t in their training set.
EpochAI have to send the questions (but not the answer key) to OpenAI in order to score the models. An overnight 2% -> 25% jump on this benchmark is a bit curious.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#136Or it's time to step back and call it what it is - very good pattern recognition. I mean, that's cool... we can get a lot of work done with pattern recognition. Most of the human race never really moves above that level of thinking in the workforce or navigating their daily life, especially if they default to various societally prescribed patterns of getting stuff done (eg. go to college or the military , find a job…
So, I am conflicted about this. If we take an example of what is considered a priori as creativity, such as story telling, LLMs can do pretty well at creating novel work. I can prompt with various parameters, plot elements, moral lessons, and get a de novo storyline, conflicts, relationships, character backstories, intrigues, and resolutions. Now, the writing style tends to be tone-deaf and poor at building tension f…
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#137One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#138Earlier quoted context omitted.
I've always assumed they removed it, because it's such a basic and fundamental part of ML training that you separate your test and train data. And yet I never see any papers even mention if/how they do this. And I wonder if they do, how do they guarantee with high reliability that their massive terabytes of data don't contain the answer.
First of all, Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly. Secondly, removing it from internet data is not 100% accurate. There are translations of the problems and solutions or references and direct match is not enough. MMLU and test set benchmarks show more resilience though in some previous research.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#139The paper includes several examples of their modified questions. There has been a substantial jump from o1-preview to o1, so I gave several samples to o1 and o1-pro ( not o1-preview), and current o1s gave the correct answer to those modified problems. SOTA changes fast.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#140One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)
Way too easy. If you think that mass and energy might be equivalent, then dimensional analysis doesn’t give you too much choice in the formula. Really, the interesting thing about E=mc^2 isn’t the formula but the assertion that mass is a form of energy and all the surrounding observations about the universe.
Also, the actual insight in 1905 was more about asking the right questions and imagining that the equivalence principle could really hold, etc. A bunch of the math predates 1905 and would be there in an AI’s training set:
https://en.m.wikipedia.org/wiki/History_of_Lorentz_transform...