Earlier quoted context omitted.
Hmm... Gemini (1.5 Flash) just aced that exact question for me: These lines celebrate the victory of the British ship HMS Victory, led by the famous Admiral Lord Nelson, in the Battle of Trafalgar in 1805. "Here's success unto the Victory": This line directly praises the ship itself, acknowledging its role in the successful battle. "and crew of noble fame": This recognizes the bravery and skill of the sailors who ser…
That's not acing the question. It's completely incorrect. What do you think the singer in "Friends in Low Places" meant in the toast he gave after crashing his ex-girlfriend's wedding? And I saw the surprise and the fear in his eyes when I took his glass of champagne and I toasted you, said "Honey, we may be through but you'll never hear me complain"
30% drop in O1-preview accuracy when Putnam problems are slightly variated
521–530 of 558 posts
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#522Earlier quoted context omitted.
10 pounds of bricks is actually heavier than 10 pounds of feathers.
Can you explain? An ounce of gold is heavier than an ounce of feathers, because the "ounce of gold" is a troy ounce, and the "ounce of feathers" is an avoirdupois ounce. But that shouldn't be true between feathers and bricks - they're both avoirdupois.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#523Earlier quoted context omitted.
o3 is able to get 25% on never seen before frontiermath problems. sure, the models do better when the answer is directly in their dataset but they’ve already surpassed the average human in novelty on held out problems
I think the problems it solved were understood to be well known undergraduate problems. https://xenaproject.wordpress.com/2024/12/22/can-ai-do-maths...
1. it suggests it’s possible that more of the problems are IMO-esque than previously thought, we don’t know how the share of solved problems is.
2. calling IMO problems “well known undergraduate problems” is a bit much
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#524I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…
I just asked Claude 3.5 Sonnet, which appears to have improved its response with CoT but there's mistakes that demonstrate the model doesn't really "understand": Q: A woman and her son are in a car accident. The woman is sadly killed. The boy is rushed to hospital. When the doctor sees the boy he says "I can't operate on this child, he is my son". How is this possible? C: Let me think about this step by step: A woman…
"The doctor is the boy’s other parent—specifically his mother, who wasn’t in the accident."
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#525Earlier quoted context omitted.
They can generalise to novel inputs. Ok often they mess it up and they're clearly better at dealing with inputs they have seen before (who isn't?), but they can still reason about things they have never seen before. Honestly if you don't believe me just go and use them. It's pretty obvious if you actually get experience with them.
Current LLMs are equivalent to tabular Markov chains (though these are too huge to realistically compute). What's the size limit when a tabular Markov chain can generalize to novel inputs?
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#526I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…
What weighs more, a 100kt aircraft carrier or a 200kt thermonuclear weapon?
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#527Earlier quoted context omitted.
I am rather saying that there is no one national drink for The Netherlands, like a Frenchman would say wine, a German/Belgian would say beer, and a Scotsman would say whisky. Note that I prompted "In the Netherlands, in terms of drinks, is there a particular spirit that represents the country?" I didn't ask which spirit is consumed the most. For example, France has been trending towards beer more and more, and within…
The question of whether there is a national drink seems to me to be entirely different than the question you asked the LLM "Prompt: In the Netherlands, in terms of drinks, is there a particular spirit that represents the country?" The question in the prompt comes off to me as a sort of qualitative determination rather than asking about pure factual information (is there an officially designated spirit). As such I don…
I was just proving the people wrong that were saying akin to that o1 was "oneshotting every question".
I completely understand from how LLMs work that they wouldn't be able to get this right. But then people shouldn't be proudly be pronouncing that o1 (or any model) is getting every question right, first time.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#528Earlier quoted context omitted.
I tried 3 times the "Which is heavier, a 10.01kg block of steel on or a 9.99kg bag of feathers?" and ChatGPT keep converting kg to pound and saying the 9.99kg is heavier.
Which model? On the paid plus tier, GPT-4o, GPT-o1, and GPT-o1mini all successfully got the 10.1. I did not try any other models. gpt-4o: https://chatgpt.com/share/67768221-6c60-8009-9988-671beadb5a... o1-mini: https://chatgpt.com/share/67768231-6490-8009-89a6-f758f0116c... o1: https://chatgpt.com/share/67768254-1280-8009-aac9-1a3b75ccb4...
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#529Earlier quoted context omitted.
That's not acing the question. It's completely incorrect. What do you think the singer in "Friends in Low Places" meant in the toast he gave after crashing his ex-girlfriend's wedding? And I saw the surprise and the fear in his eyes when I took his glass of champagne and I toasted you, said "Honey, we may be through but you'll never hear me complain"
That requires knowing the song, beyond the words provided. Would you flunk an eighth grader for getting it wrong?
But I think specifying that the singer has crashed his ex-girlfriend's wedding is already enough that you deserve to fail if your answer is "he says he's not upset, so what he means is that he's not upset". It's not any kind of leap to guess that the bride's ex-boyfriend's toast might cause a scene at a wedding - that's why the bride's ex-boyfriends are never invited.
(The question has already provided every word of the toast that appears in the song.)
See also the sidethread comment by mikeruiz, noting that o1-pro reproduces the rest of the lyrics to The Victory, but gets the question wrong anyway.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#530Earlier quoted context omitted.
If GP's hypothesis was "it fails for small variations of the input, like this one", then testing that hypothesis with that exact variation on a couple models seems fair and scientific. Testing it with more variations until one fails feels a bit like p-hacking. You'd need to engage in actual statistics to get reliable results from that, beyond "If I really try, I can make it fail". Which would be a completely differen…
I feel like I'm almost 100% certain that the smart guys at OpenAI have added many more variations of the problem to their training set since OP did his failing test, so it doesn't surprise me at all to know that this exact one now passes. In fact, in my use of o1 it's incredibly clear that it still has the same problems. It's incredibly common that the second I ask for someone even slightly outside the training set,…
May still get it wrong in more subtle ways, though. Personally, I think it'll continue to get physics wrong until someone builds it some robot arms so it can train on actually interactive physical spaces and behavior.