Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

481–490 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#481
post #49

One experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)

what if someone invented it before 1905

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#482

Earlier quoted context omitted.

Didn't they run a bunch of models on the problem set? I doubt they are hosting all those models on their own infrastructure.

1. OpenAI has confirmed it’s not in their train (unlike putnam where they have never made any such claims) 2. They don't train on API calls 3. It is funny to me that HN finds it easier to believe theories about stealing data from APIs rather than an improvement in capabilities. It would be nice if symmetric scrutiny were applied to optimistic and pessimistic claims about LLMs, but I certainly don’t feel that is the c…

The gpt 4 paper said they only excluded benchmarks by exact text matches. That means discussion of them probably doesn't get excluded.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#483
post #204

Earlier quoted context omitted.

https://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f... so far o1-mini has bodied every task people are saying LLMs can’t do in this thread

That appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be differ…

But if it doesn't get it correct every time, it doesn't "know".

By that standard humans know almost nothing.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#484

Earlier quoted context omitted.

This thread has come up before(1), but I'll continue to argue that relativistic mass is a perfectly valid concept as much as any other, and if you disagree, you'll need arguments more substantial than it just being unpopular these days. Especially if you're trying to argue people out of using a concept that they personally find useful to aid their own understanding, just because it doesn't fit your own mathematical o…

sure but i think it 1. is not very intuitive/useful to have mass that varies on the direction (which is what this implies) 2. is somewhat tautological to define a new mass m_rel = E/c^2 and say that it satisfies the equation when this is not what most people understand mass to be. most people understand photons to be massless particles. at minimum , relativistic mass should always be specified as m_rel to distinguish…

> most people understand photons to be massless particles.

I suspect that most people have no opinion at all and are generally unaware of the properties of light.

That being said, a (perfectly reflective or otherwise steady state, e.g. black body at equilibrium) box of photons has mass due to those photons. You can stick it on a scale or a pendulum and measure it. It attracts other nearby objects according to GR or, in the appropriate limit, Newtonian gravity, in accordance with the relativistic mass of those photons.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#485
sometimes o1-preview start hallucinating halfway through a good solution. it can get the intuition and the 'main' direction for a problem wrong too. but then problem solving is just a series of rephrasings and translating into different math domains is used by mahtematicians for solving problems.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#486
post #288

Earlier quoted context omitted.

> That addition is a small capability but you only need a single counterexample to disprove a theory No, that's not how this works :) You can hardcode an exception to pattern recognition for specific cases - it doesn't cease to be a pattern recognizer with exceptions being sprinkled in. The 'theory' here is that a pattern recognizer can lead to AGI. That is the theory. Someone saying 'show me proof or else I say a pa…

It's not hardcoded, reissbaker has addressed this point. I think you are misinterpreting what the argument is. The argument being made is that LLMs are mere 'stochastic parrots' and therefore cannot lead to AGI. The analogy to Russell's teapot is that someone is claiming that Russells teapot is not there because china cannot exist in the vacuum of space. You can disprove that with a single counterexample. That does n…

[deleted]

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#487
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

Llama 3.33-70B:

"To determine which is heavier, we need to compare the weights of the two objects.

The steel ingots weigh 9.99 pounds, and the fluffy cotton weighs 10.01 pounds.

Since 10.01 pounds is greater than 9.99 pounds, the bag of fluffy cotton is heavier than the bag of steel ingots, despite the cotton being "fluffy" and the steel being dense. Weight is a measure of the amount of mass in an object, and in this case, the cotton has a slightly greater mass than the steel."

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#488

Earlier quoted context omitted.

Even simpler, I asked Gemini (Flash 1.5) this variant of the question: ``` I have two bags, one can hold a pound of gold and one can hold a pound of feathers. Which bag is heavier? ``` The point here a) the question really is a bit too vague, b) if you assume that each back is made of the same material and that each bag is just big enough to hold the contents, the bag for the feathers will need to be much bigger than…

I think Gemini did better than you think with its second answer! Your original question didn't mention that the bags were made of the same material or the same density of material. The set of all possible bags that could hold 1 pound of feathers includes some thinner, weaker bags than the set of all possible bags that could hold 1 pound of gold (the gold being denser). So absent any other prior information the probab…

Saran wrap would would a pound of gold. It's just a pound in roughly a cubic inch of material. A cloth bag (like a pillow which is how features are normally contained) would definitely suffice. Remember, a pound of 800 fill down would require 8 cubic feet.

The bag for the gold doesn't need anything substantial.

So, no, the chances are strongly on the side of the feather bag being heavier.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#489

Earlier quoted context omitted.

I must be to tired as I can't find any flaw in that reasoning.

> The doctor is the boy's father. This is a classic riddle that highlights gender bias - many people assume doctors must be men, but don't initially consider that the doctor could be the father. It identified the pattern, but it didn't understand the father is a man while the mother is a woman. If people assume doctors must be men, they wouldn't consider that the doctor could be the mother. (Side note: this is an ant…

I read it as “most people don’t connect put 2 and 2 together - the doctor could be the boys father”. It is poorly worded, but not strictly wrong I think?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#490

One I just did: Q: I was heading to Aberdeen from London. On my way I passed seven wives, each wife had seven sacks, and in each sack there were seven cats and each cat had seven fish. How many were going to London? A: This riddle is a play on words, and the answer is hidden in the phrasing! You mentioned you were heading to Aberdeen from London, but you didn’t say anything about the seven wives, their sacks, cats, o…

Llama 3.33-70b-instruct:

> I was heading to Boston from New York. On my way I passed six clerks, each clerk had six socks, and in each sock there were six skinks and each skink had six sticks. How many were going to Boston?

The answer is... "I" was going to Boston! The question asks how many were going to Boston, and the answer is the person being referred to, which is "I". The rest of the information about the clerks, socks, skinks, and sticks is just a clever distraction!

[Of course, it also got your exact prompt right.]

Post reply on HN