Nah, this is just an early example of many "this is too easy it doesn't count" defensive human arguments against AI. Parallel to the "you use copilot so your code quality is terrible and you don't really even understand it so it's not maintainable" human coping we are familiar with. If there is any shred of truth to these defenses, it is temporary and will be shown false by future, more powerful AI models. Consider t…
Sorry, but a new prompt for GPT-4 is not a paper
151–160 of 193 posts
Re: Sorry, but a new prompt for GPT-4 is not a paper
#152Earlier quoted context omitted.
I'm sorry, but that's entirely ridiculous. You're mangling up a concept of burden of proof here. You can easily see this because it can be flipped around easily - you made a claim that they are being changed, even every few weeks! Should it really be on me to show that your very specific claim is false? Aside - but even if the model weights did change, that wouldn't stop research being possible. Otherwise no drug tri…
> You can easily see this because it can be flipped around easily - you made a claim that they are being changed, even every few weeks! Should it really be on me to show that your very specific claim is false? Wait a minute? The author of such a paper makes a claim about some observation that's based on the assumption that the studied model is defined in a way. I am disputing that claim since no evidence has been sho…
You are entirely within your rights to say that the authors have assumed that openai is not lying about their models. They've probably also assumed that other paper authors are not lying in their papers.
You then say however:
> GPT is a closed source/weights, proprietary product that changes every couple of weeks or so.
And when I ask for evidence of this very specific claim, you turn around and say the burden is on me to show that you're lying. That is what is butchering the concept of burden of proof.
> If your twist on this issue would be true, then I would, by definition, have to accept everything that they claim as true without any evidence.
Absolutely not.
Re: Sorry, but a new prompt for GPT-4 is not a paper
#153Earlier quoted context omitted.
>> LLM research is currently in its infancy Everything that went in to creating GPT4 is AI/science or whatever. Probing GPT4 and trying to understand and characterize it is also a very worthy thing to do - else how can it be improved upon? But if making GPT is science, I'd say this stuff is more akin to psychology ;-)
Machine psychology
Re: Sorry, but a new prompt for GPT-4 is not a paper
#154Earlier quoted context omitted.
> Compare with drug trials: Adderall only differs from regular amphetamine in the relative concentration of enantiomers, and the entire value of the drug is in the measurements. Drug trials may be expected to be somewhat reproducible. What I don't get is how it can even be called research if it cannot be expected to be reproducible at all! GPT is a closed source/weights, proprietary product that changes every couple…
You can select a static snapshot that presumably does not change, if you use the API
Re: Sorry, but a new prompt for GPT-4 is not a paper
#155Earlier quoted context omitted.
I'm sorry, but that's entirely ridiculous. You're mangling up a concept of burden of proof here. You can easily see this because it can be flipped around easily - you made a claim that they are being changed, even every few weeks! Should it really be on me to show that your very specific claim is false? Aside - but even if the model weights did change, that wouldn't stop research being possible. Otherwise no drug tri…
Apples and oranges comparison. You couldn’t get the same participants, but you could get the same drugs. If you could get identical participants, that wouldn’t be very helpful since humans are so varied. But for GPT based papers, what you’re actually testing could change without you knowing. There’s no way to know if a paper is reproducible at all. If you can’t reproduce results, is it really research, or just show a…
You can't start with a statement about clinical trials not being perfectly reproducible and that's fine, then say this.
> what you’re actually testing could change without you knowing
If people are lying about an extremely important part of their product, which they have little reason to. But then this applies to pretty much everything. Starting with the assumption that people are lying about everything and nothing is as it seems may technically make things more reproducible but it's going to require unbelievable effort for very little return.
> There’s no way to know if a paper is reproducible at all.
This is a little silly because these models are available extremely easily and at a pay-as-you-go pricing. And again, it requires an assumption that openai is lying about a specific feature of a product.
Re: Sorry, but a new prompt for GPT-4 is not a paper
#156Earlier quoted context omitted.
> Compare with drug trials: Adderall only differs from regular amphetamine in the relative concentration of enantiomers, and the entire value of the drug is in the measurements. Drug trials may be expected to be somewhat reproducible. What I don't get is how it can even be called research if it cannot be expected to be reproducible at all! GPT is a closed source/weights, proprietary product that changes every couple…
If I book telescope time and capture a supernova then no one will ever be able to reproduce my raw results because it has already happened. I don't see why OpenAI pulling old model snapshots is any different.
Re: Sorry, but a new prompt for GPT-4 is not a paper
#157Earlier quoted context omitted.
> Compare with drug trials: Adderall only differs from regular amphetamine in the relative concentration of enantiomers, and the entire value of the drug is in the measurements. Drug trials may be expected to be somewhat reproducible. What I don't get is how it can even be called research if it cannot be expected to be reproducible at all! GPT is a closed source/weights, proprietary product that changes every couple…
If I book telescope time and capture a supernova then no one will ever be able to reproduce my raw results because it has already happened. I don't see why OpenAI pulling old model snapshots is any different.
The Twitter user doesn't even reference a single specific paper, kind of doing some hand wavy broad generalizations of his worst antagonists. So who really knows what he's talking about? I can't say.
If he means papers like the ones in this search - https://arxiv.org/search/?query=step+by+step+gpt4&searchtype... - they're all kind of interesting, especially https://arxiv.org/abs/2308.06834 which is the kind of "new prompt" class he's directly attacking. It is interesting because it was written by some doctors, and it's about medicine, so it has some interdisciplinary stuff that's more interesting than the computer science stuff. So I don't even agree with the premise of what the Twitter complainer is maybe complaining about, because he doesn't name a specific paper.
Anyway, to your original point, if we're comparing the research I linked and astronomy... well, they're completely different, it is totally intellectually dishonest to compare the two. Like tell me how I use astronomy research later in product development or whatever? Maybe in building telescopes? How does observing the supernova suggest new telescopes to build in the future, without suggesting that indeed, I will be reproducing the results, because I am building a new telescope to observe another such supernova? Astronomy cares very deeply about reproducibility, a different kind of reproducibility than these papers, but maybe more the same in interesting ways than the non-difference you're talking about. I'm not an astronomer, but if you want to play the insight porn game, I'd give these people a benefit of the doubt.
Re: Sorry, but a new prompt for GPT-4 is not a paper
#158Some of the most interesting papers published this year ("Automatic Multi-Step Reasoning and Tool-Use") compare prompt strategies across a variety of tasks. The results are fascinating, findings are applicable and invite further research in the area of "prompt selection" or "tool selection."
Re: Sorry, but a new prompt for GPT-4 is not a paper
#159I am sorry but what can ChatGPT do that a couple of minutes of googling couldn’t solved? Write half hearted essays that all contain the same phrase?
Re: Sorry, but a new prompt for GPT-4 is not a paper
#160To bring some data to a sour grapes fight: https://paperswithcode.com/sota/code-generation-on-humaneval For code generation, GPT4 is getting beat by the small prompt library LATS wrapped around GPT3.5. Given the recent release of MagicCoder / Instruct-OSS, that means a small prompt library + a small 7B model you can self-host beats the much fancier GPT4. Similar to when simple NNs destroyed a decade of Bayesian model…
The link you shared doesn’t quite reflect this. Omitting other models…
LATS (gpt-4): 94.4 Reflexion (gpt-4): 91.0 gpt-4: 86.6 … LATS (gpt-3.5): 83.8 … zero-shot (gpt-4): 67.0 zero-shot (gpt-3.5): 48.1
I’m not quite sure how to translate leaderboards like these into actual utility, but it certainly feels like “good enough” is only going to get more accessible and I agree with what I think is your broader point - more sophisticated techniques will make small, affordable, self-hostable models viable in their own right.
I’m optimistic we’re on a path where further improvement isn’t totally dependent on just throwing money at more parameters.