Earlier quoted context omitted.
you sure? i just asked o1-mini ( not 4o mini ) 5 times in a row (new chats obviously) and it got it right every time perhaps you stumbled on a rarer case but reading the logs you posted this sounds more like a 4o model than an o1 because it’s doing its thinking in the chat itself plus the procedure you described would probably get you 4o-mini
> just asked o1-mini (not 4o mini) 5 times in a row (new chats obviously) and it got it right every time Could you try playing with the exact numbers and/or substances?
30% drop in O1-preview accuracy when Putnam problems are slightly variated
241–250 of 558 posts
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#242Earlier quoted context omitted.
You mean train on pre-1939 data and predict how WWII would go?
Right. If it were trained through August 1939, how much prompting would be necessary to get it to predict aspects of WWII.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#243Earlier quoted context omitted.
https://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f... so far o1-mini has bodied every task people are saying LLMs can’t do in this thread
That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…
The default model I see on chatgpt.com is GPT 4o-mini, which is not o1-mini.
OpenAI describes GPT 4o-mini as "Our fast, affordable small model for focused tasks" and o1/o1-mini as "Reasoning models that excel at complex, multi-step tasks".
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#244Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#245Earlier quoted context omitted.
Why does AI have to be smarter than the collective of hummanity in order to be considered intelligent? It seems like we keep raising the bar on what intelligence means ¯\_(ツ)_/¯
A machine that synthesizes all human knowledge really ought to know more than an individual in terms of intellect. An entity with all of human intellect prior to 1905 does not need to be as intelligent as a human to make discoveries that mere humans with limited intellect made. Why lower the bar?
Because of the chance of misundertanding. Failing at acknowledging artificial general intelligence standing right next to us.
An incredible risk to take in alignment.
Perfect memory doesn't equal to perfect knowledge, nor perfect understanding of everything you can know. In fact, a human can be "intelligent" with some of his own memories and/or knowledge, and - more commmonly - a complete "fool" with most of the rest of his internal memories.
That said, is not a bit less generally intelligent for that.
Supose it exists a human with unlimited memory, it retains every information touching any sense. At some point, he/she will probably understand LOTs of stuff, but it's simple to demonstrate he/she can't be actually proficient in everything: you have read how do an eye repairment surgery, but have not received/experimented the training,hence you could have shaky hands, and you won't be able to apply the precise know-how about the surgery, even if you remember a step-by-step procedure, even knowing all possible alternatives in different/changing scenarios during the surgery, you simply can't hold well the tools to go anywhere close to success.
But you still would be generally intelligent. Way more than most humans with normal memory.
If we'd have TODAY an AI with the same parameters as the human with perfect memory, it will be most certainly closely examined and determined to be not a general artificial intelligence.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#246Earlier quoted context omitted.
That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…
While this may be true, it's a very common problem that people who want to demonstrate how bad a model is fail to provide a direct link or simply state the name of the model.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#247Earlier quoted context omitted.
sounds like you had a chemistry education. relativistic mass is IMO very much not a useful way of thinking about this and it is sort of tautologically true that E = m_relativistic because “relativistic mass” is just taking the concept of energy and renaming it “mass”
This is all sort of silly IMO. The equation, like basically all equations, needs context. What’s E? What’s m? If E is the total energy of the system and m is the mass (inertial or gravitational? how far past 1905 do you want to go?), then there isn’t a correction. If m is rest mass and E is total energy, then I would call it flat-out wrong, not merely approximate. After all, a decent theory really ought to reproduce…
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#248I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…
GPT 4:
The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. The type of material doesn’t affect the weight comparison; it’s purely a matter of which bag weighs more on the scale.
GPT 4o:
The 10.01-pound bag of fluffy cotton is heavier. Weight is independent of the material, so the bag of cotton’s 10.01 pounds outweighs the steel ingots’ 9.99 pounds.
GPT o1:
Since both weights are measured on the same scale (pounds), the 10.01-pound bag of cotton is heavier than the 9.99-pound bag of steel, despite steel being denser. The key is simply that 10.01 pounds exceeds 9.99 pounds—density doesn’t affect the total weight in this comparison.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#249Earlier quoted context omitted.
> That addition is a small capability but you only need a single counterexample to disprove a theory No, that's not how this works :) You can hardcode an exception to pattern recognition for specific cases - it doesn't cease to be a pattern recognizer with exceptions being sprinkled in. The 'theory' here is that a pattern recognizer can lead to AGI. That is the theory. Someone saying 'show me proof or else I say a pa…
GPT-4o doesn't have hardcoded math exceptions. If you would like something verifiable, since we don't have the source code to GPT-4o, consider that Qwen 2.5 72b can also add large integers, and we do have the source code and weights to run it... And it's just a neural net. There isn't secret "hardcode an exception to pattern recognition" in there that parses out numbers and adds them. The neural net simply learned to…
Is the claim then that LLMs are pattern recognizers but also more?
It just seems to me and I guess many others that the thing it is primarily good at is being a better google search.
Is there something big that I and presumably many others are missing and if so, what is it?
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#25099% of studies claiming some out of distribution failure of an LLM uses a model already made irrelevant by SOTA. These kinds of studies, with long throughputs and review periods, are not the best format to make salient points given the speed at which the SOTA horizon progresses