Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

541–550 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#541

Earlier quoted context omitted.

That's not acing the question. It's completely incorrect. What do you think the singer in "Friends in Low Places" meant in the toast he gave after crashing his ex-girlfriend's wedding? And I saw the surprise and the fear in his eyes when I took his glass of champagne and I toasted you, said "Honey, we may be through but you'll never hear me complain"

Sounds like you just proved ted_dunning isn't sentient.

Well, I proved that he's happy to express an opinion on whether an answer to a question is correct regardless of whether he knows anything about the question. I wouldn't trust advice from him or expect his work output to stand up to scrutiny.

Sentience isn't really a related concept.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#542

Earlier quoted context omitted.

Well, the rest of the song helps, in that it specifies that (1) the toast upset the wedding, and (2) the singer responded to that by insulting "you", which is presumably one or more of the bride, the groom, and the guests. But I think specifying that the singer has crashed his ex-girlfriend's wedding is already enough that you deserve to fail if your answer is "he says he's not upset, so what he means is that he's no…

I was referring to the original query, of course, as any entity capable of reasoning could have figured out.

Hmm. Is there anything in my comment above that might address that point of view?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#543

Earlier quoted context omitted.

Lots of other websites are more appropriate for meme jokes.

Like I said.

Your two word comment was ambiguous. I interpreted it as something like "People are downvoting you because they have no sense of humor".

There are other websites where two and three word comments are better received.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#544

Earlier quoted context omitted.

Like I said.

Your two word comment was ambiguous. I interpreted it as something like "People are downvoting you because they have no sense of humor" . There are other websites where two and three word comments are better received.

Mea culpa.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#545

Earlier quoted context omitted.

i'd prefer an easily verifiable question rather than one where we can always go "no that's not what they really meant" but someone else with o1-mini quota can respond

It's not a difficult or tricky question.

i think it's a bit tricky, the surface meaning is extremely praiseworthy and some portion of readers might interpret as someone who has praise for Admiral Nelson but hates the press gangs.

of course, it is a sardonic, implicit critique of Admiral Nelson/the victory, etc. but i do think it is a bit subtle.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#546

Earlier quoted context omitted.

I was referring to the original query, of course, as any entity capable of reasoning could have figured out.

Hmm. Is there anything in my comment above that might address that point of view?

Nah, intermittent failures are apparently enough to provide evidence that an entire class of entities is incapable of reason. So I think we've figured this one out...

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#547
post #507

Earlier quoted context omitted.

> smart people generally earn more. > there's been a widespread belief that being wealthy is the proof of superiority. Both of these are assumptions though, and working in the reverse order. Its one thing to expect that intelligence will lead to higher value outcomes and entirely different to expect that higher value outcomes prove intelligence. It seems reasonable that higher intelligence, combined with the incentiv…

I think we're in agreement? I'm saying their measure in this case is no worse than any other, but not that it's a fundamental truth. All the other things — chess, Jeopardy, composing music, painting, maths, languages, passing medical or law degrees — they're also all things which were considered signs of intelligence until AI got good at them. Goodhart's law keeps tripping us up on the concept of intelligence.

> I think we're in agreement? I'm saying their measure in this case is no worse than any other, but not that it's a fundamental truth.

Maybe we are? I think I lost the thread a bit here.

> chess, Jeopardy, composing music, painting, maths, languages, passing medical or law degrees

That's interesting, I would have still chalked skill in those areas as a sign of intelligence and didn't realize most people wouldn't once AI (or ML) could do it. To me an AI/LLM/ML being good at those is at least a sign that they have gotten good at mimicking intelligence if nothing else, and a sign that we really are getting out over our skis risking these tools without knowing how they really work.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#548

Earlier quoted context omitted.

No, a baby is pre-trained. We know from linguistics that there is a natural language grammar template all humans follow. This template is intrinsic to our biology and is encoded and not learned through observation.

A baby has a template but so does an LLM. The better comparison to the templating is all the labor that went into making the LLM, not how long the GPUs run. Template versus template, or specific training versus specific training. Those comparisons make a lot more sense than going criss-cross.

The template is what makes the training process so short for humans. We need minimal data and we can run off of that.

Training is both longer and less effective for the LLM because there is no template.

To give an example suppose it takes just one picture for a human to recognize a dog and it takes 1 million pictures for a ML model to do the same. What I’m saying is that it’s like this because humans come preprogrammed with application specific wetware to do the learning and recognition as a generic operation. That’s why it’s so quick. For AI we are doing it as a one shot operation on something that is not application specific. The training takes longer because of this and is less effective.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#549
post #368

Earlier quoted context omitted.

I’m not really sure what you’re trying to say here - that LLMs don’t work like human brains? We don’t need to conduct any analyses to know that LLMs don’t “know” anything in the way humans “know” things because we know how LLMs work. That doesn’t mean that LLMs aren’t incredibly powerful; it may not even mean that they aren’t a route to AGI.

>We don’t need to conduct any analyses to know that LLMs don’t “know” anything in the way humans “know” things because we know how LLMs work. People, including around HN, constantly argue (or at least phrase their arguments) as if they believed that LLMs do, in fact, possess such "knowledge". This very comment chain exists because people are trying to defend against a trivial example refuting the point - as if there…

> I don't accept your definition of "intelligence" if you think that makes sense. Systems must be able to know things in the way that humans (or at least living creatures) do, because intelligence is exactly the ability to acquire such knowledge.

According to whom? There is certainly no single definition of intelligence, but most people who have studied it (psychologists, in the main) view intelligence as a descriptor of the capabilities of a system - e.g., it can solve problems, it can answer questions correctly, etc. (This is why we call some computer systems "artificially" intelligent.) It seems pretty clear that you're confusing intelligence with the internal processes of a system (e.g. mind, consciousness - "knowing things in the way that humans do").

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#550

Earlier quoted context omitted.

A baby has a template but so does an LLM. The better comparison to the templating is all the labor that went into making the LLM, not how long the GPUs run. Template versus template, or specific training versus specific training. Those comparisons make a lot more sense than going criss-cross.

The template is what makes the training process so short for humans. We need minimal data and we can run off of that. Training is both longer and less effective for the LLM because there is no template. To give an example suppose it takes just one picture for a human to recognize a dog and it takes 1 million pictures for a ML model to do the same. What I’m saying is that it’s like this because humans come preprogramm…

I disagree that an LLM has no template, but this is getting away from the point.

Did you look at the post I was replying to? You're talking about LLMs being slower, while that post was impressed by LLMs being "faster".

They're posing it as if LLMs recreate the same templating during their training time, and my core point is disagreeing with that. The two should not be compared so directly.

Post reply on HN