Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

441–450 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#441
post #430

Earlier quoted context omitted.

Ideas are kinda there, details (like names) are wrong wrong. Feel like I should make this task my personal ASI benchmark lol :D

o1-pro just answered (it often takes a few minutes). Not sure if this is any better (is it?) but at least it is starting to admit when it isn't sure about something, rather than just spewing BS confidently. That seems like a genuine win. ---------snip-------- From the handwriting and context, this appears to be a mid-19th-century petition submitted to the Orphans’ Court of Baltimore County. Below is a point‐by‐point…

She said it's around 60% there, but not helpful for her specifically as her area of research is on the families of the slave trade, and so in the document, that is actually the only thing that really matters to her. Names and places (she spends A LOT of time tracking transfer of slaves through states) - I guess I should be honest that I'm goalpost moving a bit from my original original post, it can work through some 18th english century text, but generally struggles where it matters, the details.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#442
post #441

Earlier quoted context omitted.

o1-pro just answered (it often takes a few minutes). Not sure if this is any better (is it?) but at least it is starting to admit when it isn't sure about something, rather than just spewing BS confidently. That seems like a genuine win. ---------snip-------- From the handwriting and context, this appears to be a mid-19th-century petition submitted to the Orphans’ Court of Baltimore County. Below is a point‐by‐point…

She said it's around 60% there, but not helpful for her specifically as her area of research is on the families of the slave trade, and so in the document, that is actually the only thing that really matters to her. Names and places (she spends A LOT of time tracking transfer of slaves through states) - I guess I should be honest that I'm goalpost moving a bit from my original original post, it can work through some…

i’m curious what the ground truth actually names in the doc are

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#443
post #430

Earlier quoted context omitted.

Ideas are kinda there, details (like names) are wrong wrong. Feel like I should make this task my personal ASI benchmark lol :D

o1-pro just answered (it often takes a few minutes). Not sure if this is any better (is it?) but at least it is starting to admit when it isn't sure about something, rather than just spewing BS confidently. That seems like a genuine win. ---------snip-------- From the handwriting and context, this appears to be a mid-19th-century petition submitted to the Orphans’ Court of Baltimore County. Below is a point‐by‐point…

o1-pro takes images? had no idea

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#444
post #152

Earlier quoted context omitted.

Everyone has a very different idea of what the word "intelligence" means; this definition has got the advantage that, unlike when various different AI became superhuman at arithmetic, symbolic logic, chess, jeopardy, go, poker, number of languages it could communicate in fluently, etc., it's tied to tasks people will continuously pay literally tens of trillions of dollars each year for because they want those tasks d…

This definition alone might be fine enough if the word "intelligence" wasn't already widely used outside of AI research. It is though, and the idea that intelligence is measured solely through economic value is a very, very strange approach. Try applying that definition to humans and you pretty quickly run into issues, both moral and practical. It also invalidates basically anything we've done over centuries consider…

> This definition alone might be fine enough if the word "intelligence" wasn't already widely used outside of AI research. It is though, and the idea that intelligence is measured solely through economic value is a very, very strange approach.

The response from @s1mplicissimus' on my previous comment is asking about "common usage" definitions of intelligence, and this is (IMO unfortunately) one of the many "common usage" definitions: smart people generally earn more.

I don't like "commmon sense" anything (or even similar phrases), because I keep seeing the phrase used as a thought-terminating cliché — but one thing it does do, is make it not "a very, very strange approach".

Wrong, that happens a lot for common language, but it can't really be strange.

> Try applying that definition to humans and you pretty quickly run into issues, both moral and practical.

Yes. But one also runs into issues with all definitions of it that I've encountered.

> It also invalidates basically anything we've done over centuries considering what intelligence is and how to measure it.

Sadly, not so. Even before we had IQ tests (for all their flaws), there's been a widespread belief that being wealthy is the proof of superiority. In theory, in a meritocracy, it might have been, but in practice not only to we not live in a meritocracy (to claim we do would deny both inheritance and luck), but also the measures of intelligence that society has are… well, I was thinking about Paul Merton and Boris Johnson the other day, so I'll link to the blog post: https://benwheatley.github.io/blog/2024/04/07-12.47.14.html

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#445
post #441

Earlier quoted context omitted.

She said it's around 60% there, but not helpful for her specifically as her area of research is on the families of the slave trade, and so in the document, that is actually the only thing that really matters to her. Names and places (she spends A LOT of time tracking transfer of slaves through states) - I guess I should be honest that I'm goalpost moving a bit from my original original post, it can work through some…

i’m curious what the ground truth actually names in the doc are

Caroline Timmis is the boy mother, she is married to Henry Jenkins but was married previously, the boy is James Timmis.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#446

Earlier quoted context omitted.

ChatGPT Plus user here. The following are all fresh sessions and first answers, no fishing. GPT 4: The 10.01-pound bag of fluffy cotton is heavier than the 9.99-pound bag of steel ingots. The type of material doesn’t affect the weight comparison; it’s purely a matter of which bag weighs more on the scale. GPT 4o: The 10.01-pound bag of fluffy cotton is heavier. Weight is independent of the material, so the bag of cot…

they've likely read this thread and adjusted their pre-filter to give the correct answer

[deleted]

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#447
post #197

I remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went al…

Even simpler, I asked Gemini (Flash 1.5) this variant of the question: ``` I have two bags, one can hold a pound of gold and one can hold a pound of feathers. Which bag is heavier? ``` The point here a) the question really is a bit too vague, b) if you assume that each back is made of the same material and that each bag is just big enough to hold the contents, the bag for the feathers will need to be much bigger than…

I think Gemini did better than you think with its second answer! Your original question didn't mention that the bags were made of the same material or the same density of material. The set of all possible bags that could hold 1 pound of feathers includes some thinner, weaker bags than the set of all possible bags that could hold 1 pound of gold (the gold being denser). So absent any other prior information the probability is greater than 50% that the gold-bag would be heavier than the feather-bag on that basis.

One could go further into the linguistic nuance of saying "this can hold one pound of [substance]", which often implies that that's its maximum carrying capacity; this would actually make the "trick question" answer all the more correct, as a bag that is on the cusp of ripping when holding one pound of feathers would almost certainly rip when holding one pound of (much denser) gold.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#448
post #445

Earlier quoted context omitted.

i’m curious what the ground truth actually names in the doc are

Caroline Timmis is the boy mother, she is married to Henry Jenkins but was married previously, the boy is James Timmis.

Looking more closely at the image, it seems that the strikeouts are what are confusing the model. It sees "Henry" and disregards "James" written above it. In a few places, the strikeouts almost look more like underlining.

Gotta be an insanely-challenging task for a program that wasn't even written with handwriting recognition in mind.

Other than the proper names, are any major details wrong?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#449
post #445

Earlier quoted context omitted.

Caroline Timmis is the boy mother, she is married to Henry Jenkins but was married previously, the boy is James Timmis.

Looking more closely at the image, it seems that the strikeouts are what are confusing the model. It sees "Henry" and disregards "James" written above it. In a few places, the strikeouts almost look more like underlining. Gotta be an insanely-challenging task for a program that wasn't even written with handwriting recognition in mind. Other than the proper names, are any major details wrong?

Dr. Jang said: "I've not transcribed this work yet it's future work. I can send it once it's transcribed, however from skimming the various chatgpt texts you sent me, the people are generally wrong, their relationships are inconsistent between all the text you sent me, these issues are why I do not go to chatgpt to help me with my research. The document generally is indeed about requesting the courts oversee the apprenticeship work. Minors could work, however, when minors are used it needed to be overseen by the courts, this document is the court overseeing the work"

She provided this as one she just got done working with: https://s.h4x.club/z8u9xmv7 (John King Esq. but try giving it to an LLM)

I will also happily again admit a bit of goal post moving on my part. I was probably a little to harsh on it (maybe because I'm used to her and her history geeks talking about how they don't work well for their research).

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#450

Earlier quoted context omitted.

Not hardcoded, I think it's just likely that those problems exist in its training data in some form

Yea, people have a really hard time dealing with data leakage especially on data sets as large as LLMs need. Basically if something appeared online or was transmitted over the wire should no longer be eligible to evaluate on. D. Sculley had a great talk at NeurIPS 2024 (same conference this paper was in) titled Empirical Rigor at Scale – or, How Not to Fool Yourself Basically no one knows how to properly evaluate LLM…

No, an absolute massive amount of people do. In fact they have been doing exactly as you recommend, because as you note, it's obvious and required for a basic proper evaluation.
Post reply on HN