Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

81–90 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#82
post #49

One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)

This reminds me of a similar idea I recently heard in podcast with Adam Brown. I'm unsure whether it is his original notion. The idea being, that if we can create AI that can derive special relativity (1905) from pre-Einstein books and papers then we have reached the next game-changing milestone in the advancement of artificial reasoning.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#83
post #52

Or it's time to step back and call it what it is - very good pattern recognition. I mean, that's cool... we can get a lot of work done with pattern recognition. Most of the human race never really moves above that level of thinking in the workforce or navigating their daily life, especially if they default to various societally prescribed patterns of getting stuff done (eg. go to college or the military , find a job…

So, I am conflicted about this. If we take an example of what is considered a priori as creativity, such as story telling, LLMs can do pretty well at creating novel work. I can prompt with various parameters, plot elements, moral lessons, and get a de novo storyline, conflicts, relationships, character backstories, intrigues, and resolutions. Now, the writing style tends to be tone-deaf and poor at building tension f…

To add to this pondering: we are discussing the state today, right now. We could assume this is as good as it's ever gonna get, and all attempts to overcome some current plateau are futile, but I wouldn't bet on it. There is a solid chance that 8th grade level writer will turn into a post-grad writer before long.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#84
post #49

One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)

There is a reason why they won't do it. They are selling a narrative. There is a lot of money to be made here with this narrative and proving that artificial intelligence is NOT intelligent won't help sell that narrative.

The goal is to make it intelligent, by which OpenAI in particular explicitly mean "economically useful", not simply to be shiny.

Passing tests is well known to be much easier than having deep understanding, even in humans. They openly ask for tests like this, not that they could possibly prevent them if they wanted to.

There's scammers trying what you say of course, and I'm sure we've all seen some management initiatives or job advertisements for some like that, but I don't get that impression from OpenAI or Anthropic, definitely not from Apple or Facebook (LeCun in particular seems to deny models will ever do what they actually do a few months later). Overstated claims from Microsoft perhaps (I'm unimpressed with the Phi models I can run locally, GitHub's copilot has a reputation problem but I've not tried it myself), and Musk definitely (I have yet to see someone who takes Musk at face value about Optimus).

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#85
post #63

They are highly effective pattern matchers. You change the pattern, it won't work. I don't remember who, but most likely @tszzl (roon), commented on x that they still trained the traditional way, and there is no test time compute (TTC) or Montecarlo Tree search (like Alpha Go) in o1 or o3. If that is true, then it's still predicting the next word based on it's training data. Likely to follow the most probable path -…

I recently watched two interviews with OpenAI researchers where they describe that the breakthrough of o-series (unlike GPT series) is to focus on test time compute as they are designed to “think” more specifically to avoid pattern matching. Noam Brown https://youtu.be/OoL8K_AFqkw?si=ocIS0YDXLvaX9Xb6&t=195 and Mark Chen https://youtu.be/kO192K7_FaQ?si=moWiwYChj65osLGy

[deleted]

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#86
post #77

Earlier quoted context omitted.

> Good students are immune to variations I don't believe that. I'd put some good money that if an excellent student is given an exact question from a previous year, they'll do better (faster & more accurate) on it, than when they're given a variation of it.

I don't think you are betting on the same thing the parent comment is talking about. The assumptions aren't the same to begin with.

What's the difference between benefitting from seeing previous problems and being worse off when not having a previous problem to go from?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#87
post #49

One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)

Finally, a true application of E=mc^2+AI

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#89
post #6

Earlier quoted context omitted.

Seems a bit picky. If the bot has seen the exact problem before it's not really doing anything more than recall to solve it.

20 years ago in grad school we were doing a very early iteration of this where we built Markov chains with Shakespeare's plays and wanted to produce a plausibly "Shakespearian" clause given a single word to start and a bearish professor said "the more plausible it gets the more I worry people might forget plausibility is all that it promises". (There was also a much earlier piece of software that would generate semi-…

I think your prof’s worries came true on a massive scale

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#90
post #47
post #41

Earlier quoted context omitted.

You're saying that humans perform worse on problems that are slightly different than previously published forms of the same problem? To be clear we are only talking about changing variable names and constants here.

That is the principle behind the game 'Simon says'

'Simon says' is about reaction time and pressure.
Post reply on HN