Live data from Hacker News

30% drop in O1-preview accuracy when Putnam problems are slightly variated

openreview.net

151–160 of 558 posts

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#151
post #84

Earlier quoted context omitted.

There is a reason why they won't do it. They are selling a narrative. There is a lot of money to be made here with this narrative and proving that artificial intelligence is NOT intelligent won't help sell that narrative.

The goal is to make it intelligent, by which OpenAI in particular explicitly mean "economically useful", not simply to be shiny. Passing tests is well known to be much easier than having deep understanding, even in humans. They openly ask for tests like this, not that they could possibly prevent them if they wanted to. There's scammers trying what you say of course, and I'm sure we've all seen some management initiat…

I haven't seen "intelligent" used as "economically useful" anywhere outside the AI hype bubble. The most charitable interpretation I can think of is lack of understanding of the common usage of the word, the most realistic one is intentionally muddying terminology so one cannot be called a liar. Are LLMs helpful tools for some tasks like rough translations, voice2text etc? Sure. Does it resemble what humans call intelligence? I'd yet have to see an example of that. The suggested experiment is a great idea and would sway my opinion drastically (given all the training data, model config, prompts & answers are public and reproducible of course, we don't want any chance of marketing BS to taint the results, do we). I'll be honest though, I'm not going to hold my breath for that experiment to succeed with the LLM technology...

edit: lol downvoted for calling out shilling i guess

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#152
post #84

Earlier quoted context omitted.

The goal is to make it intelligent, by which OpenAI in particular explicitly mean "economically useful", not simply to be shiny. Passing tests is well known to be much easier than having deep understanding, even in humans. They openly ask for tests like this, not that they could possibly prevent them if they wanted to. There's scammers trying what you say of course, and I'm sure we've all seen some management initiat…

> The goal is to make it intelligent, by which OpenAI in particular explicitly mean "economically useful", not simply to be shiny I never understood why this definition isn't a huge red flag for most people. The idea of boiling what intelligence is down to economic value is terrible, and inaccurate, in my opinion.

Everyone has a very different idea of what the word "intelligence" means; this definition has got the advantage that, unlike when various different AI became superhuman at arithmetic, symbolic logic, chess, jeopardy, go, poker, number of languages it could communicate in fluently, etc., it's tied to tasks people will continuously pay literally tens of trillions of dollars each year for because they want those tasks done.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#153
post #126

The researcher's answer to their variant of "Year: 2016 ID: A1" in the appendix is wrong. The solution (sum of 1,2,5,6,9,10,13,14, ...) has an alternating pattern, so has to be two piecewise interleaved polynomials, which cannot be expressed as a single polyomial. Their answer works for k=1,2, but not k=3. https://openreview.net/pdf?id=YXnwlZe0yf This does not give me confidence in the results of their paper.

You are correct. Their answer is instead the sum of the first k terms of 1, 2, 6, 10, 14, 18, ..., for positive k.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#154

Earlier quoted context omitted.

Didn't they run a bunch of models on the problem set? I doubt they are hosting all those models on their own infrastructure.

1. OpenAI has confirmed it’s not in their train (unlike putnam where they have never made any such claims) 2. They don't train on API calls 3. It is funny to me that HN finds it easier to believe theories about stealing data from APIs rather than an improvement in capabilities. It would be nice if symmetric scrutiny were applied to optimistic and pessimistic claims about LLMs, but I certainly don’t feel that is the c…

The modern state of training is to try to use everything they can get their hands on. Even if there are privileged channels that are guaranteed not to be used as training data, mentioning the problems on ancillary channels (say emailing another colleague to discuss the problem) can still create a risk of leakage because nobody making the decision to include the data is aware that stuff that should be excluded is in that data set. And as we've seen from decades of cybersecurity, people are absolute shit at the necessary operational security to avoid mentioning stuff on ancillary channels!

Given that performance is known to drop considerably on these kinds of tests when novel problems are tried, and given the ease with which these problems could leak into the training set somehow, it's not unreasonable to be suspicious of a sudden jump in performance as merely a sign that the problems made it into the training set rather than being true performance improvements in LLMs.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#155
post #110

Earlier quoted context omitted.

And what prompt would you give that does have novel input.

If I was me, I would start by giving a collection of LLMs the patent, ask half "why is this patent novel" and half "why is this patent not novel" and see what happens. I use this method of "debugging" my thinking (not code), might be a starting point here? Not sure.

LLMs are already good at summarizing the claims - patents all explain why they’re novel - so it would be a waste to ask them, especially if you reserve half the LLMs in your set for this question. Asking why a patent is not novel is a great question, but the problem with asking why they are not novel is it has to know all other patents (including very recently filed patents) and it has to be correct, which LLMs are not at all good at yet (plus they still tend to hallucinate confidently). This is a great test for LLM accuracy if you know the right answer already, and not a good test for patent validity.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#156

Earlier quoted context omitted.

1. OpenAI has confirmed it’s not in their train (unlike putnam where they have never made any such claims) 2. They don't train on API calls 3. It is funny to me that HN finds it easier to believe theories about stealing data from APIs rather than an improvement in capabilities. It would be nice if symmetric scrutiny were applied to optimistic and pessimistic claims about LLMs, but I certainly don’t feel that is the c…

Easier to believe or not, thinking that it's not a reasonable possibility is also funny.

Do you also think they somehow stole the codeforces problems before they were even written or you are willing to believe the #175 global rank there?

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#157
I wouldn’t be surprised if similar will be found concerning the ARC challenge and it is why I still maintain my own private LLM challenges to gauge current capabilities. Course, I have little illusion that these are fully private, but it is better than fully public tests.

Even the most straight forward, logical, easily reasoned ones stump all LLMs I have access to, which is why I am so skeptical concerning emergence, reasoning and all this hype around “AGI”…

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#158

Earlier quoted context omitted.

Why does AI have to be smarter than the collective of hummanity in order to be considered intelligent? It seems like we keep raising the bar on what intelligence means ¯\_(ツ)_/¯

A machine that synthesizes all human knowledge really ought to know more than an individual in terms of intellect. An entity with all of human intellect prior to 1905 does not need to be as intelligent as a human to make discoveries that mere humans with limited intellect made. Why lower the bar?

The heightening of the bar is an attempt to deny that milestones were surpassed and to claim that LLMs are not intelligent.

We had a threshold for intelligence. An LLM blew past it and people refuse to believe that we passed a critical milestone in creating AI. Everyone still thinks all an LLM does is regurgitate things.

But a technical threshold for intelligence cannot have any leeway for what people want to believe. They don’t want to define an LLM as intelligent even if it meets the Turing test technical definition of intelligence so they change the technical definition.

And then they keep doing this without realizing and trivializing it. I believe humanity will develop an entity smarter than humans but it will not be an agi because people keep unconsciously moving the goal posts and changing definitions without realizing it.

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#159

Earlier quoted context omitted.

It certainly feels like certain patterns are hardcoded special cases, particularly to do with math. "Solve (1503+5171)*(9494-4823)" reliably gets the correct answer from ChatGPT "Write a poem about the solution to (1503+5171)*(9494-4823)" hallucinates an incorrect answer though That suggests to me that they've papered over the models inability to do basic math, but it's a hack that doesn't generalize beyond the simpl…

“a poem about” reads to me at least like the solution need not be in the answer; maybe something like “a poem that includes the answer in the last stanza”

[deleted]

Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated

#160
post #137

Earlier quoted context omitted.

I had a similar thought but about asking the LLM to predict “future” major historical events. How much prompting would it take to predict wars, etc.?

You mean train on pre-1939 data and predict how WWII would go?

Right. If it were trained through August 1939, how much prompting would be necessary to get it to predict aspects of WWII.
Post reply on HN