Earlier quoted context omitted.
https://chatgpt.com/share/67756c29-111c-8002-b203-14c07ed1e6... I got a very different answer: A 10.01-pound bag of fluffy cotton is heavier than a 9.99-pound bag of steel ingots because 10.01 pounds is greater than 9.99 pounds. The material doesn't matter in this case; weight is the deciding factor. What model returned your answer?
You also didn't ask the question correctly.
30% drop in O1-preview accuracy when Putnam problems are slightly variated
231–240 of 558 posts
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#232Earlier quoted context omitted.
A machine that synthesizes all human knowledge really ought to know more than an individual in terms of intellect. An entity with all of human intellect prior to 1905 does not need to be as intelligent as a human to make discoveries that mere humans with limited intellect made. Why lower the bar?
The heightening of the bar is an attempt to deny that milestones were surpassed and to claim that LLMs are not intelligent. We had a threshold for intelligence. An LLM blew past it and people refuse to believe that we passed a critical milestone in creating AI. Everyone still thinks all an LLM does is regurgitate things. But a technical threshold for intelligence cannot have any leeway for what people want to believe…
> We had a threshold for intelligence.
We’ve had many. Computers have surpassed several barriers considered to require intelligence such as arithmetic, guided search like chess computers, etc etc. the Turing test was a good benchmark because of how foreign and strange it was. It’s somewhat true we’re moving the goalposts. But the reason is not stubbornness, but rather that we can’t properly define and subcategorize what reason and intelligence really is. The difficulty to measure something does not mean it doesn’t exist or isn’t important.
Feel free to call it intelligence. But the limitations are staggering, given the advantages LLMs have over humans. They have been trained on all written knowledge that no human could ever come close to. And they still have not come up with anything conceptually novel, such as a new idea or theorem that is genuinely useful. Many people suspect that pattern matching is not the only thing required for intelligent independent thought. Whatever that is!
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#233Earlier quoted context omitted.
Okay, then what about elite level codeforces performance? Those problems weren’t even constructed until after the model was made. The real problem with all of these theories is most of these benchmarks were constructed after their training dataset cutoff points. A sudden performance improvement on a new model release is not suspicious. Any model release that is much better than a previous one is going to be a “sudden…
Can you give an example of one of these problems that 'wasn't even constructed until after the model was made'? I'd like to see if it's truly novel and unique, the first problem of its type ever construed by mankind, or if it's similar to existing problems.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#234Earlier quoted context omitted.
https://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f... so far o1-mini has bodied every task people are saying LLMs can’t do in this thread
This happens literally every time. Someone always says "ChatGPT can't do this!", but then when someone actually runs the example, chatGPT gets it right. Now what the OP is going to do next is proceed to move goalposts and say like "but umm I just asked chatgpt this, so clearly they modified the code in realtime to get the answer right"
I mean, if I had OpenAI’s resources I’d have a team tasked with monitoring social to debug trending fuck-ups. (Before that: add compute time to frequently-asked novel queries.)
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#235Earlier quoted context omitted.
That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…
you sure? i just asked o1-mini ( not 4o mini ) 5 times in a row (new chats obviously) and it got it right every time perhaps you stumbled on a rarer case but reading the logs you posted this sounds more like a 4o model than an o1 because it’s doing its thinking in the chat itself plus the procedure you described would probably get you 4o-mini
Could you try playing with the exact numbers and/or substances?
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#236Earlier quoted context omitted.
I didn't downvote it, but short comments are a very big risk. People may misinterpret it, or think it's crackpot theory or a joke and then downvote. When in doubt, add more info, like: But the complete equation is E=sqrt(m^2c^4+p^2) that is reduced to E=mc^2 when the momentum p is 0. More info in https://en.wikipedia.org/wiki/Mass%E2%80%93energy_equivalenc...
The next section of the wikipedia link discusses the low speed approximation, where sqrt(m^2c^4+(pc)^2) ≈ mc^2 + 1/2 mv^2. Calling E=mc^2 an "approximation" is technically correct. It's the 0th order approximation. That's just pointlessly confusing. A better word choice would be "a special case".
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#237Earlier quoted context omitted.
What I learnt is that there is a rest mass and a relativistic mass. The m in your formula is the rest mass. But when you use the relativistic mass E=mc² still holds. And for the rest mass I always used m_0 to make clear what it is.
sounds like you had a chemistry education. relativistic mass is IMO very much not a useful way of thinking about this and it is sort of tautologically true that E = m_relativistic because “relativistic mass” is just taking the concept of energy and renaming it “mass”
IMO, when people get excited about E=mc^2, it’s in contexts like noticing that atoms have rest masses that are generally somewhat below the mass of a proton or neutron times the number of protons and neutrons in the atom, and that the mass difference is the binding energy of the nucleus, and you can do nuclear reactions and convert between mass and energy! And then E=mc^2 is apparently exactly true, or at least true to an excellent degree, even though the energies involved are extremely large and Newtonian mechanics can’t even come close to accounting for what’s going on.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#238Earlier quoted context omitted.
The heightening of the bar is an attempt to deny that milestones were surpassed and to claim that LLMs are not intelligent. We had a threshold for intelligence. An LLM blew past it and people refuse to believe that we passed a critical milestone in creating AI. Everyone still thinks all an LLM does is regurgitate things. But a technical threshold for intelligence cannot have any leeway for what people want to believe…
Since we know an LLM does indeed simply regurgitate data, having it pass a "test for intelligence" simply means that either the test didn't actually test intelligence, or that intelligence can be defined as simply regurgitating data.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#239Earlier quoted context omitted.
https://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f... so far o1-mini has bodied every task people are saying LLMs can’t do in this thread
That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…
Does this proves he is not an intelligent being?
Is he stupid?
This he had a lapse? Would we judge his intelligence for that?
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#240Earlier quoted context omitted.
I don't see any reason to assume they removed it unless they're very explicit about it. Model publishers have an extremely strong vested interest in beating benchmarks and I expect them to teach to the test if they can get away with it.
As usual, once a metric becomes a target, it stops being useful.