Performance of these LLMs on real life tasks feels very much like students last-minute cramming for Asian style exams. The ability to perfectly regurgitate, while no concept of meaning.
o3 is able to get 25% on never seen before frontiermath problems. sure, the models do better when the answer is directly in their dataset but they’ve already surpassed the average human in novelty on held out problems
o3 is able to get 25% on never seen before frontiermath problems. sure, the models do better when the answer is directly in their dataset but they’ve already surpassed the average human in novelty on held out problems
> never seen before frontiermath problems How do you know that?
Because that is the whole conceit of how frontiermath is constructed
The paper includes several examples of their modified questions. There has been a substantial jump from o1-preview to o1, so I gave several samples to o1 and o1-pro ( not o1-preview), and current o1s gave the correct answer to those modified problems. SOTA changes fast.
LLM boosters are so tiresome. You hardly did a rigorous evaluation, the set has been public since October and could have easily been added to the training data.
I've always assumed they removed it, because it's such a basic and fundamental part of ML training that you separate your test and train data. And yet I never see any papers even mention if/how they do this. And I wonder if they do, how do they guarantee with high reliability that their massive terabytes of data don't contain the answer.
First of all, Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly. Secondly, removing it from internet data is not 100% accurate. There are translations of the problems and solutions or references and direct match is not enough. MMLU and test set benchmarks show more resilience though in some previous research.
OpenAI is extremely cagey about what's in their test data set generally, but absent more specific info, they're widely assumed to be grabbing whatever they can. (Notably including copyrighted information used without explicit authorization -- I'll take no position on legal issues in the New York Times's lawsuit against OpenAI, but at the very least, getting their models to regurgitate NYT articles verbatim demonstrates pretty clearly that those articles are in the training set.)
One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)
This reminds me of a similar idea I recently heard in podcast with Adam Brown. I'm unsure whether it is his original notion. The idea being, that if we can create AI that can derive special relativity (1905) from pre-Einstein books and papers then we have reached the next game-changing milestone in the advancement of artificial reasoning.
Right, hadn't listened to that one, thanks for the tip!
o3 is able to get 25% on never seen before frontiermath problems. sure, the models do better when the answer is directly in their dataset but they’ve already surpassed the average human in novelty on held out problems
The average human did zero studying on representative problems. LLMs did a lot .
One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)
But is there even enough pre-1905 data to create models that say hello world reliably? The terabytes of training data required for decent LLMs does not exist. I’d guess there may only be gigabytes worth.
My wife is an 18th century American history professor. LLMs have very very clearly not been trained on 18th century English, they cannot really read it well, and they don't understand much from that period outside of very textbook stuff, anything nuanced or niche is totally missing. I've tried for over a year now, regularly, to help her use LLMs in her research, but as she very amusingly often says "your computers are useless at my work!!!!"
Performance of these LLMs on real life tasks feels very much like students last-minute cramming for Asian style exams. The ability to perfectly regurgitate, while no concept of meaning.
o3 is able to get 25% on never seen before frontiermath problems. sure, the models do better when the answer is directly in their dataset but they’ve already surpassed the average human in novelty on held out problems