Live data from Hacker News

Sorry, but a new prompt for GPT-4 is not a paper

twitter.com

11–20 of 193 posts

Re: Sorry, but a new prompt for GPT-4 is not a paper

#11
post #7

If you do enough measurements on that new prompt then I don't see why this shouldn't be a paper. People overestimate the value of "grand developments", and underestimate the value of actually knowing - in this case actually knowing how well something works, even if it is as simple as a prompt. Compare with drug trials: Adderall only differs from regular amphetamine in the relative concentration of enantiomers, and th…

> People overestimate the value of "grand developments", and underestimate the value of actually knowing - in this case actually knowing how well something works, even if it is as simple as a prompt.

I think this depends a lot on the "culture" of the subject area. For example in mathematics, it is common that only new results that have been thoroughly worked through are typically "publish-worthy".

Re: Sorry, but a new prompt for GPT-4 is not a paper

#12
post #7

If you do enough measurements on that new prompt then I don't see why this shouldn't be a paper. People overestimate the value of "grand developments", and underestimate the value of actually knowing - in this case actually knowing how well something works, even if it is as simple as a prompt. Compare with drug trials: Adderall only differs from regular amphetamine in the relative concentration of enantiomers, and th…

> People overestimate the value of "grand developments", and underestimate the value of actually knowing - in this case actually knowing how well something works, even if it is as simple as a prompt. I think this depends a lot on the "culture" of the subject area. For example in mathematics, it is common that only new results that have been thoroughly worked through are typically "publish-worthy".

Wouldn’t the “thoroughly worked through” part be analogous to extensive measurements of a prompt?

Re: Sorry, but a new prompt for GPT-4 is not a paper

#13

Real science is reserved for those with real expertise! As the self-anointed gatekeeper of real science I decree that other peoples’ work fails to meet the minimum standard I have set for real science! Mind you not the work other actors in the scientific community publish and accept among their peers - they are not real scientists and their work is trivial. For shame!

The real shame is that even HN has fallen into the trap of missing obvious and funny sarcasm unless it is clearly labeled.

Re: Sorry, but a new prompt for GPT-4 is not a paper

#14

Attacking the participants in a systemic shift is 100% useless as it doesn't target the culprit. In programming we have a similar phenomenon, that StackOverflow-driven (and I guess now GPT-driven) juniors have overtaken the industry and displaced serious talent. Because sufficient amounts of quantity always beats quality, even if the end result is inferior, this is caused by market dynamics, which operate on much cru…

This whole thing reminds me a bit of playing Fallout 4.

There will be good data, the pre-AI enshitification data. The stuff from before the war.

And then... the data after. Tainted by the entropy, and lack of utility of AI.

Alas, this means in some senses, human progress will slow and stop in the tech field if we aren't careful and preserve ways to create pre-AI data. But the cost of it is so high in comparison to post... I'm not sold it will be worth it.

Re: Sorry, but a new prompt for GPT-4 is not a paper

#15
post #7

If you do enough measurements on that new prompt then I don't see why this shouldn't be a paper. People overestimate the value of "grand developments", and underestimate the value of actually knowing - in this case actually knowing how well something works, even if it is as simple as a prompt. Compare with drug trials: Adderall only differs from regular amphetamine in the relative concentration of enantiomers, and th…

> Compare with drug trials: Adderall only differs from regular amphetamine in the relative concentration of enantiomers, and the entire value of the drug is in the measurements.

Drug trials may be expected to be somewhat reproducible.

What I don't get is how it can even be called research if it cannot be expected to be reproducible at all!

GPT is a closed source/weights, proprietary product that changes every couple of weeks or so. How can you expect a prompt to do the same for a reasonable length of time for the research to be even rudimentarily reproducible? And if it's not reproducible, what is it actually worth? I don't think much. Could as well have been a fault in the research setup or a fake.

Re: Sorry, but a new prompt for GPT-4 is not a paper

#17

Attacking the participants in a systemic shift is 100% useless as it doesn't target the culprit. In programming we have a similar phenomenon, that StackOverflow-driven (and I guess now GPT-driven) juniors have overtaken the industry and displaced serious talent. Because sufficient amounts of quantity always beats quality, even if the end result is inferior, this is caused by market dynamics, which operate on much cru…

> Because sufficient amounts of quantity always beats quality

Hegel (as echoed by Marx): "merely quantitative differences beyond a certain point pass into qualitative changes" ( https://www.pnas.org/doi/10.1073/pnas.240462397 )

I always found that an interesting observation, whatever you think of the rest of their works.

Re: Sorry, but a new prompt for GPT-4 is not a paper

#18
post #7

If you do enough measurements on that new prompt then I don't see why this shouldn't be a paper. People overestimate the value of "grand developments", and underestimate the value of actually knowing - in this case actually knowing how well something works, even if it is as simple as a prompt. Compare with drug trials: Adderall only differs from regular amphetamine in the relative concentration of enantiomers, and th…

> Compare with drug trials: Adderall only differs from regular amphetamine in the relative concentration of enantiomers, and the entire value of the drug is in the measurements. Drug trials may be expected to be somewhat reproducible. What I don't get is how it can even be called research if it cannot be expected to be reproducible at all! GPT is a closed source/weights, proprietary product that changes every couple…

> GPT is a closed source/weights, proprietary product that changes every couple of weeks or so.

Do you have any evidence that the weights for versioned models are being changed without notifications?

Re: Sorry, but a new prompt for GPT-4 is not a paper

#19
post #7

If you do enough measurements on that new prompt then I don't see why this shouldn't be a paper. People overestimate the value of "grand developments", and underestimate the value of actually knowing - in this case actually knowing how well something works, even if it is as simple as a prompt. Compare with drug trials: Adderall only differs from regular amphetamine in the relative concentration of enantiomers, and th…

> Compare with drug trials: Adderall only differs from regular amphetamine in the relative concentration of enantiomers, and the entire value of the drug is in the measurements. Drug trials may be expected to be somewhat reproducible. What I don't get is how it can even be called research if it cannot be expected to be reproducible at all! GPT is a closed source/weights, proprietary product that changes every couple…

You can select a static snapshot that presumably does not change, if you use the API

Re: Sorry, but a new prompt for GPT-4 is not a paper

#20
post #18

Earlier quoted context omitted.

> Compare with drug trials: Adderall only differs from regular amphetamine in the relative concentration of enantiomers, and the entire value of the drug is in the measurements. Drug trials may be expected to be somewhat reproducible. What I don't get is how it can even be called research if it cannot be expected to be reproducible at all! GPT is a closed source/weights, proprietary product that changes every couple…

> GPT is a closed source/weights, proprietary product that changes every couple of weeks or so. Do you have any evidence that the weights for versioned models are being changed without notifications?

> Do you have any evidence that the weights for versioned models are being changed without notifications?

I think in a real scientific process, it's upon those who claim that they are not to provide the evidence.

Post reply on HN