Live data from Hacker News

Sorry, but a new prompt for GPT-4 is not a paper

twitter.com

161–170 of 193 posts

Re: Sorry, but a new prompt for GPT-4 is not a paper

#161

To bring some data to a sour grapes fight: https://paperswithcode.com/sota/code-generation-on-humaneval For code generation, GPT4 is getting beat by the small prompt library LATS wrapped around GPT3.5. Given the recent release of MagicCoder / Instruct-OSS, that means a small prompt library + a small 7B model you can self-host beats the much fancier GPT4. Similar to when simple NNs destroyed a decade of Bayesian model…

This right here. I feel like the focus on just throwing more GPU at the problem is a mistake many of these companies are making at the moment. The real breakthroughs will come when we figure out how to use the current models and compute power more efficiently. If it’s prompt engineering that leads to this breakthrough, so be it.

Re: Sorry, but a new prompt for GPT-4 is not a paper

#162
post #70

Why not? A paper is not necessarily scientific nor a breakthrough. In my view, a paper is written and documented communication that's usually approved by peers in the field. Also a blunt observation in nature can be noteworthy. However, we don't see such papers anymore as these fields have matured. Just go back in the history of your field and you will find trivial papers.

In the medical field, letters and case studies often document observations that may not be groundbreaking. However, scientific journals typically feature content that contributes to existing knowledge, making it somewhat novel. Consequently, presenting a set of POST parameters as an arXiv paper could be perceived as undermining the integrity of the entire preprint service.

Re: Sorry, but a new prompt for GPT-4 is not a paper

#163
post #68

It is a _paper_, but it's not science, since GTP-4 is closed source and thus not reproducible in a lab. If OpenAI disappears tomorrow, papers are GTP-4 will likely be of little to no value, which is another tell of a non-scientific exploration. (note: not all explorations are scientific, and that is great! Science is just one of many tools for exploring lived reality.)

That’s like saying a biologist studying an endangered species isn’t doing science because the animal could disappear tomorrow. The permanence of a subject has no bearing on whether it is science or not. The idea that science has to happen in a lab is of course absurd as well.

In the case of an endangered species a biologist would still have access to take samples from it and inspect it. Science doesn't have to happen in a lab but it's questionable to call something science when it involves hitting a black box endpoint which can change the underlying models and behaviors at a whim.

Re: Sorry, but a new prompt for GPT-4 is not a paper

#164
post #127
post #73

Earlier quoted context omitted.

I feel like the author of this tweet wasn’t saying step-by-step isn’t worthy, he was saying that non-reproducible results are not science. He emphasizes this twice in that tweet: > one experiment on one data set with seed picking is not worthy reporting > Additionally, we all need to understand this is just one good empirical result, now we need to make it useful…

Is it non-reproducible? Also results which reproducibility can be measured and appears stable is perfectly good science. I dislike when people throw statements like that.

I have no idea. The author of that tweet seems to imply that the results aren’t reproducible. I was just commenting to point out that the author’s intent may have been different from what the grandparent comment was saying.

Re: Sorry, but a new prompt for GPT-4 is not a paper

#165

Earlier quoted context omitted.

Why stop at chemistry? Chemistry is fundamentally quantum electrodynamics applied to huge ensembles of particles. QED is very well understood and gives the best predictions we have to date of any scientific theory. How come we don’t entirely understand biology then?

> Why stop at chemistry? Chemistry is fundamentally quantum electrodynamics applied to huge ensembles of particles. Chemistry is indeed applied QED ;) (and you don't need massive numbers of particles to have very complex chemistry) > How come we don’t entirely understand biology then? We understand some of the basics (even QED is not reality). That understanding comes from bottom-up studies of biochemistry, but most…

I will ignore your patronizing remarks beyond acknowledging them here, in order to promote civil discourse.

I think you have missed my point by focusing on biology as an extremely complex field.e, it was my mistake to use it as an example in the first place. We don’t need to go that far;

sure, llms did not spawn on their own. They are a result of thousands of years of progress in countless fields of science and engineering. Like any modern invention, essentially.

Here I remember to make sure we are on the same page on what we’re discussing - as I understand, whether “prompt engineering” can be considered an engineering/science practice. Personally I haven’t considered this enough to form an opinion but your argument does not sound convincing to me;

I guess your idea of what llms represent matters here. The way I see it, in some abstract sense we are as society exploring a current peak - in compute $ or flops and performance on certain tasks - of a rather large but also narrow family of functions. By focusing our attention on functions composed of ones we understood how to effectively find parameters for, we were able to build at this point rather complicated processes for finding parameters for the compositions.

Yes, the components are understood, at various levels of rigor, but the thing produced is not yet sufficiently understood. Partly out of cost to reproduce such research, and partly due to complexity of the system, a driver for the cost.

The fact that “prompt engineering” as a practice and that companies supposedly base their business model on secret prompts is a testament, for me, to the fact they are not well understood. A well understood system you design has a well understood interface.

Now, I haven’t noticed a specific post OP was criticizing so i take it his remarks were general. He seems to thinks that some research is not worth publishing. I tend to agree that I would like research to be of high quality, but that is subjective. Is it novel? is it true?

Now, progress will be progress and im sure current architectures will change and models will get larger. And it may be that a few giants are the only one running models large enough to require prompt engineering. Or we may find a way to have those models understand us better than a human ever could. Doubtful. And post singularity anyway, by definition.

In either case yes, probably temporary profession. But in case open research will continue in those directions as well, there will be need for people to figure out ways to communicate effectively with these. You dismiss them as testers.

However, progress in science and engineering is often driven by data where theory is lacking and I’m not aware of the existence of deep theory as of yet. eg something that would predict how well a certain architecture would perform. Engineering ahead of theory, driven by $).

As in physics that we both mentioned, knowing the component part does not automatically grant you understanding of the whole. knowing everything there is to know about the relevant physical interaction, protein folding was a tough problem that AFAIR has had a lot of success with tools from the field. Square in the realm of physics even, and we can’t give good predictions without testing (computationally).

If someone tested some folding algorithm and visually inspected results, then found a trick how to consistently improve on the result in some subcase of proteins. Would that be worthy of publishing? if yes, why is this different? if not, why not?

Re: Sorry, but a new prompt for GPT-4 is not a paper

#166
post #42

Nah, this is just an early example of many "this is too easy it doesn't count" defensive human arguments against AI. Parallel to the "you use copilot so your code quality is terrible and you don't really even understand it so it's not maintainable" human coping we are familiar with. If there is any shred of truth to these defenses, it is temporary and will be shown false by future, more powerful AI models. Consider t…

No prompt will cause an LLM to rapidly improve itself, much less into an AGI. Prompts don't cause permanent change in the LLM, only differences in output.

You're talking about how GPT functions in 2023. I am discussing such a point where when LLM outputs become valuable LLM modifications.

AI recursing on itself progressing toward an AGI.

Re: Sorry, but a new prompt for GPT-4 is not a paper

#167
post #152

Earlier quoted context omitted.

> You can easily see this because it can be flipped around easily - you made a claim that they are being changed, even every few weeks! Should it really be on me to show that your very specific claim is false? Wait a minute? The author of such a paper makes a claim about some observation that's based on the assumption that the studied model is defined in a way. I am disputing that claim since no evidence has been sho…

> I am disputing that claim You are entirely within your rights to say that the authors have assumed that openai is not lying about their models. They've probably also assumed that other paper authors are not lying in their papers. You then say however: > GPT is a closed source/weights, proprietary product that changes every couple of weeks or so. And when I ask for evidence of this very specific claim, you turn arou…

[deleted]

Re: Sorry, but a new prompt for GPT-4 is not a paper

#168
post #13

Real science is reserved for those with real expertise! As the self-anointed gatekeeper of real science I decree that other peoples’ work fails to meet the minimum standard I have set for real science! Mind you not the work other actors in the scientific community publish and accept among their peers - they are not real scientists and their work is trivial. For shame!

The real shame is that even HN has fallen into the trap of missing obvious and funny sarcasm unless it is clearly labeled.

Sarcasm is often spoken with a sarcastic inflection. It doesn't translate well to text, regardless of the community.

Re: Sorry, but a new prompt for GPT-4 is not a paper

#169

Earlier quoted context omitted.

you realize nobody understands WHY or HOW these models work under the hood right? it's akin to evolution - we understand the process - that part is simple. But the output/organisms we have to investigate how they work.

> you realize nobody understands WHY or HOW these models work under the hood right? Of course we understand how they work, we built them! There is no mystery in their mechanisms, we know the number of neurons, their connectivity, everything from the weights to the activation functions. This is not a mystery, this is several decades of technical developments. > it's akin to evolution - we understand the process - that…

We designed the process. We didn't design the models - the models were "designed" based on the features of a massive dataset and massive number of iterations.

Even if you understand evolution - you still don't understand how the human body or mind works. That needs to be investigated and discovered.

In the same way, you understanding how these models were trained doesn't help you understand how the models work. That needs to be investigated and discovered.

Re: Sorry, but a new prompt for GPT-4 is not a paper

#170

Earlier quoted context omitted.

You can select a static snapshot that presumably does not change, if you use the API

> You can select a static snapshot that presumably does not change, if you use the API Sorry, I won't blindly believe a company who are cynical enough to call themselves "OpenAI", then publish a commercial closed source/weights model for profit. Evidence that they do not change without notice or it didn't happen. Better even, provide the source and weights for research purposes. These models could be pulled at every…

Yeah, here it comes. In these conversations you don’t need to ask very many “why”s before it just turns out that the antagonist (you) has an axe to grind about OpenAI, and has added that the their misplaced sense of expertise with regard to the typical standards of proof in academic publications.
Post reply on HN