Live data from Hacker News

Sorry, but a new prompt for GPT-4 is not a paper

twitter.com

91–100 of 193 posts

Re: Sorry, but a new prompt for GPT-4 is not a paper

#92

Earlier quoted context omitted.

If I book telescope time and capture a supernova then no one will ever be able to reproduce my raw results because it has already happened. I don't see why OpenAI pulling old model snapshots is any different.

> If I book telescope time and capture a supernova then no one will ever be able to reproduce my raw results because it has already happened. I don't see why OpenAI pulling old model snapshots is any different. That's why you capture multiple of them and verify your data statistically?

And ideally if someone is proposing new prompting techniques they should test it across both the most capable models (which are unfortunately proprietary) and the best open models.

The problem is that what works on small LLMs does not necessarily scale to larger ones. See page 35 of [1] for example. A researcher only using the models of a few years ago (where the open models had [1] https://arxiv.org/pdf/2308.03296.pdf

Re: Sorry, but a new prompt for GPT-4 is not a paper

#93
post #83

Excuse me? Step by step wasn't paper-worthy? Hard disagree. LLM research is currently in its infancy, because they are no older than a few years old. And a research field in its infancy is bound to have a few noteworthy "no sh*t, Sherlock" papers that would be obvious from hindsight. The fact is, LLMs are a higher-order construct in machine learning, much like a fish is higher-order than a simple cellular colony. Low…

Sorry, ignoramus here: Which paper is “Step by step”?

I think it's a reference to the "discovery" that if you ask GPT-4 to answer your query "step by step", it'll actually offer a better response than otherwise.

Re: Sorry, but a new prompt for GPT-4 is not a paper

#94
post #61

Earlier quoted context omitted.

What is incorrect in this reference? You have not proposed any counterarguments. Also if you need just more fresh data - how do you propose to interpret the result of the Rosenhan's experiment?

Dont put people inside mental asylums when they are not ill?

There is no evidence of existing at least one defined illness in psy* fields. For example, let me tell you that a person X fell ill with schizophrenia. What do you know about X or X's brain?

Re: Sorry, but a new prompt for GPT-4 is not a paper

#95

Excuse me? Step by step wasn't paper-worthy? Hard disagree. LLM research is currently in its infancy, because they are no older than a few years old. And a research field in its infancy is bound to have a few noteworthy "no sh*t, Sherlock" papers that would be obvious from hindsight. The fact is, LLMs are a higher-order construct in machine learning, much like a fish is higher-order than a simple cellular colony. Low…

On a related note, there is this recent tweet purportedly showing that "offering to give a tip to ChatGPT" improves performance (or at the very least resulted in longer responses, which might not be a good proxy for performance) https://twitter.com/voooooogel/status/1730726744314069190

Re: Sorry, but a new prompt for GPT-4 is not a paper

#96
post #73

Excuse me? Step by step wasn't paper-worthy? Hard disagree. LLM research is currently in its infancy, because they are no older than a few years old. And a research field in its infancy is bound to have a few noteworthy "no sh*t, Sherlock" papers that would be obvious from hindsight. The fact is, LLMs are a higher-order construct in machine learning, much like a fish is higher-order than a simple cellular colony. Low…

I feel like the author of this tweet wasn’t saying step-by-step isn’t worthy, he was saying that non-reproducible results are not science. He emphasizes this twice in that tweet: > one experiment on one data set with seed picking is not worthy reporting > Additionally, we all need to understand this is just one good empirical result, now we need to make it useful…

Exactly, and I tend to agree with him. I argued some time ago here that a paper should take some time to try to explain why its results are happening, at least from a reasonable hypothesis (people didn't seem to agree). An experiment (even a simple one) starts from a null hypothesis and tries to disprove it. However, most of what we see coming out of "scientific" papers is basically just engineering, I guess?: we put all of these things together in some way (out of pure guess and/or preference bias) and these results happened. We don't know why, good luck figuring it out. Here is one example where it works (don't ask where it doesn't; we intentionally kept those out).

And while I obviously value very much the engineering advances we have seen, the science is still lacking, because not enough people are trying to understand why these things are happening. Although engineering advances are important and valuable, I don't understand exactly why people try so hard to call themselves scientists if they are basically skipping the scientific process entirely.

Re: Sorry, but a new prompt for GPT-4 is not a paper

#97
post #12

Earlier quoted context omitted.

> People overestimate the value of "grand developments", and underestimate the value of actually knowing - in this case actually knowing how well something works, even if it is as simple as a prompt. I think this depends a lot on the "culture" of the subject area. For example in mathematics, it is common that only new results that have been thoroughly worked through are typically "publish-worthy".

Wouldn’t the “thoroughly worked through” part be analogous to extensive measurements of a prompt?

Let me put it this way: you can expect that a typical good math paper means working on the problem for, I would say, half a year (often much longer). I have a feeling that most papers that involve extensive measurements of prompts do not involve 1/2 to 1 year of careful

- hypothesis building

- experimental design

- doing experiments

- analyzing the experimental results

- doing new experiments

- analyzing in which sense the collected data support the hypothesis or not

- ...

work.

Re: Sorry, but a new prompt for GPT-4 is not a paper

#99

Earlier quoted context omitted.

>> LLM research is currently in its infancy Everything that went in to creating GPT4 is AI/science or whatever. Probing GPT4 and trying to understand and characterize it is also a very worthy thing to do - else how can it be improved upon? But if making GPT is science, I'd say this stuff is more akin to psychology ;-)

Machine psychology

Asimov was way ahead of you:

https://en.wikipedia.org/wiki/Robopsychology

Re: Sorry, but a new prompt for GPT-4 is not a paper

#100
post #7

If you do enough measurements on that new prompt then I don't see why this shouldn't be a paper. People overestimate the value of "grand developments", and underestimate the value of actually knowing - in this case actually knowing how well something works, even if it is as simple as a prompt. Compare with drug trials: Adderall only differs from regular amphetamine in the relative concentration of enantiomers, and th…

Agreed. People publish papers on algorithms all the time, imagine saying "Sorry, but new C++ is not a paper". There is a ton of space to be explored wrt prompts.

If you do the rigor on why something really is interesting, publish it.

Post reply on HN