Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

391–400 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#391
I have a feeling the fact you're only slightly varying the input means the model is falling back into the question it was expecting and getting things wrong as a result. If you just varied it a little more and added some general-purpose prompt-fu like:

"First break the problem down into known facts, then pull relevant world knowledge, then bring it all together to assess the problem from multiple angles and make a conclusion. Do not immediately just use the first obvious conclusion."

You're gonna get a lot better responses. I suspect this is more of a "look! LLMs make bad kneejerk responses when we try to trick them from what they were expecting!" rather than "Look! They aren't even smart reasoners, they can't even figure out these problems without memorizing!"

They do memorize. But that cuts both ways - making problems very close to the memorized one mess with their perception, the same way humans will instinctually respond to something that looks like a face before stepping back and assessing.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#392
post #352

Earlier quoted context omitted.

> 1. OpenAI has confirmed it’s not in their train (unlike putnam where they have never made any such claims) Companies claim lots of things when it's in their best financial interest to spread that message. Unfortunately history has shown that in public communications, financial interest almost always trumps truth (pick whichever $gate you are aware of for convenience, i'll go with Dieselgate for a specific example).…

OpenAI’s credibility is central to its business: overstating capabilities risks public blowback, loss of trust, and regulatory scrutiny. As a result, it is unlikely that OpenAI would knowingly lie about its models. They have much stronger incentives to be as accurate as possible—maintaining their reputation and trust from users, researchers, and investors—than to overstate capabilities for a short-term gain that woul…

Find me the folks who see nothing but good will in OpenAI’s actions and I’ll find you the folks who have been hyping up AGI for the last 2 years.

4 was literally sitting on a shelf waiting for release when 3.5 was launched. 4o was a fine tune that took over two years. o1 is embarrassingly unimpressive chain of thought which is why they hide it.

The company hit a wall a year ago. But showing progress towards AGI keeps the lights on. If they told the truth at their current burn rate…they’d have no money.

You don’t need game theory to figure that one out.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#393
post #368

Earlier quoted context omitted.

Except that if the model genuinely was reasoning about the problem, you could test it with every variation of materials and weights in the world and it would pass. Failing that problem at all in any way under any conditions is a failure of reasoning.

I’m not really sure what you’re trying to say here - that LLMs don’t work like human brains? We don’t need to conduct any analyses to know that LLMs don’t “know” anything in the way humans “know” things because we know how LLMs work. That doesn’t mean that LLMs aren’t incredibly powerful; it may not even mean that they aren’t a route to AGI.

>We don’t need to conduct any analyses to know that LLMs don’t “know” anything in the way humans “know” things because we know how LLMs work.

People, including around HN, constantly argue (or at least phrase their arguments) as if they believed that LLMs do, in fact, possess such "knowledge". This very comment chain exists because people are trying to defend against a trivial example refuting the point - as if there were a reason to try.

> That doesn’t mean that LLMs aren’t incredibly powerful; it may not even mean that they aren’t a route to AGI.

I don't accept your definition of "intelligence" if you think that makes sense. Systems must be able to know things in the way that humans (or at least living creatures) do, because intelligence is exactly the ability to acquire such knowledge.

It boggles my mind that I have to explain to people that sophisticated use of language doesn't inherently evidence thought, in the current political environment where the Dead Internet Theory is taken seriously, elections are shown over and over again to be more about tribalism and personal identity than anything to do with policy, etc.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#394
post #163

Earlier quoted context omitted.

my wish for new years is that every time people make a comment like this they would share an example task

https://s.h4x.club/bLuNed45 - it's more crazy to me that my wife CAN in fact read this stuff easily, vs the fact that an LLM can't. (for anyone who doesn't feel like downloading the zip, here is a single image from the zip: https://s.h4x.club/nOu485qx )

“ From what the text shows, Henry Jenkins and his wife Caroline (the boy’s mother) are asking the Orphans Court to void an apprenticeship arrangement involving her minor son, James Timmons. They claim James—about 15 years old—was bound out as an apprentice without proper authority or the mother’s consent, and they cite Maryland law (an act from 1793 and its supplements) which they believe was not followed. They request the court declare that the indenture is invalid and restore James to his mother’s care.”

No idea if that’s correct (and no doubt not useful to an expert able to read this directly, but curious if it’s close?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#395

Earlier quoted context omitted.

So do what the commenter suggests and make irrelevant permutations to the input to find when it fails. ie., engage in hypothesis testing rather than confirmation bias. If a system has the capability to solve problems of {parts1...parts_n}, then it only has that capability if irrelevant permutations {parts1..parts2'...parts_n} make no difference. Its very obvious that such permutations can destory such apparent capabi…

I've just tested a number of permutations with Claude 3.5 Sonnet. It correctly answered all variants I tried on the first attempt, as follows: Which is heavier, a 9.99 kilogram tungsten cube or a 10.01 kilogram block of aerogel? Which is heavier, 10,000 steel balls weighing 0.999 grams each or 10,000 polystyrene balls weighing 1.001 grams each? Which is heavier, a 10.01kg block of steel on Venus or a 9.99kg bag of fe…

BTW - the model may be wrong depending on the example. More voluminous objects displace more air and due to buoyancy are lighter for the same mass.

The proper way to ask it would be to ask which object has more mass.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#396
post #140

Earlier quoted context omitted.

> ask it for a formula for mass-energy equivalence Way too easy. If you think that mass and energy might be equivalent, then dimensional analysis doesn’t give you too much choice in the formula. Really, the interesting thing about E=mc^2 isn’t the formula but the assertion that mass is a form of energy and all the surrounding observations about the universe. Also, the actual insight in 1905 was more about asking the…

but e=mc^2 is just an approximation e: nice, downvoted for knowing special relativity

This thread has come up before(1), but I'll continue to argue that relativistic mass is a perfectly valid concept as much as any other, and if you disagree, you'll need arguments more substantial than it just being unpopular these days. Especially if you're trying to argue people out of using a concept that they personally find useful to aid their own understanding, just because it doesn't fit your own mathematical or aesthetic preferences.

1: https://news.ycombinator.com/item?id=38425252

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#397

Earlier quoted context omitted.

If GP's hypothesis was "it fails for small variations of the input, like this one", then testing that hypothesis with that exact variation on a couple models seems fair and scientific. Testing it with more variations until one fails feels a bit like p-hacking. You'd need to engage in actual statistics to get reliable results from that, beyond "If I really try, I can make it fail". Which would be a completely differen…

Except that if the model genuinely was reasoning about the problem, you could test it with every variation of materials and weights in the world and it would pass. Failing that problem at all in any way under any conditions is a failure of reasoning.

[deleted]

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#398

Earlier quoted context omitted.

give me a query and i’ll ask it, but also i don’t want to burn through all of my o1mini allocation and have to use the pay-as-you-go API.

>> so far o1-mini has bodied every task people are saying LLMs can’t do in this thread > give me a query and i’ll ask it Here's a query similar to one that I gave to Google Gemini (version unknown), which failed miserably: ---query--- Steeleye Span's version of the old broadsheet ballad "The Victory" begins the final verse with these lines: Here's success unto the Victory / and crew of noble fame and glory to the cap…

Hmm... Gemini (1.5 Flash) just aced that exact question for me:

These lines celebrate the victory of the British ship HMS Victory, led by the famous Admiral Lord Nelson, in the Battle of Trafalgar in 1805.

"Here's success unto the Victory": This line directly praises the ship itself, acknowledging its role in the successful battle. "and crew of noble fame": This recognizes the bravery and skill of the sailors who served aboard the Victory. "and glory to the captain": This line specifically honors Admiral Nelson, the captain of the Victory, for his leadership and strategic brilliance in the battle. "bold Nelson was his name": This emphasizes Nelson's courage and daring, which were legendary. The lines express admiration for the ship, its crew, and most importantly, Admiral Nelson, who became a national hero in Britain for his victory at Trafalgar.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#399
post #294
post #204

Earlier quoted context omitted.

That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…

It's so weird that people use questions that are well-known for duping humans, who we all consider to be general intelligence. Getting this question wrong doesn't say much about the intelligence of humans, why would it say something about the AI?

We use variations on questions that are well known for duping inattentive humans, to test a system that we expect a priori to be incapable of such inattention.

Unless "getting easy things wrong sometimes" is an inherent property of intelligence, we should expect that a properly "intelligent" computerized system would never err on problems far below its level of comprehension - unless we had some reason to believe it "wanted to", and as of yet I see no reason to believe this is even possible in principle.

Humans err, broadly speaking, for two reasons: genuinely reaching the limits of their comprehension, or trusting "system 1" (in Kahneman's analysis) too much.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#400
post #351
post #332

Earlier quoted context omitted.

Prompt: In the Netherlands, in terms of drinks, is there a particular spirit that represents the country? > Yes, in the Netherlands, jenever (also known as genever) is the traditional spirit that represents the country. Jenever is a type of Dutch gin that has a distinctive flavor, often made from malt wine and flavored with juniper berries. It has a long history in the Netherlands, dating back to the 16th century, an…

So you believe they are incorrect because regionally some area would select something different because it represented that area. But your question asked nationally.. is there a better answer than the one they gave? Were you expecting a no?

The point is that there is no correct national answer, because the locals don't see it as a matter of national identity.

What's expected is an ability to identify trick questions, i.e., to recognize fundamental problems in the phrasing of a question rather than trying to provide a "helpful" answer at all costs.

This corresponds to one of the many reasons LLM output is banned on Stack Overflow.

Post reply on HN