Sorry, but a new prompt for GPT-4 is not a paper
141–150 of 193 posts
Re: Sorry, but a new prompt for GPT-4 is not a paper
#142Excuse me? Step by step wasn't paper-worthy? Hard disagree. LLM research is currently in its infancy, because they are no older than a few years old. And a research field in its infancy is bound to have a few noteworthy "no sh*t, Sherlock" papers that would be obvious from hindsight. The fact is, LLMs are a higher-order construct in machine learning, much like a fish is higher-order than a simple cellular colony. Low…
I have a question For people in the field! Would GPT only with pretraining and no fine tuning be able to behave better when prompt is "let’s think step by step"? Or this prompt only worked because a fine tuning dataset containing many "let’s think step by step" prompt was used?
Re: Sorry, but a new prompt for GPT-4 is not a paper
#143Earlier quoted context omitted.
I'm not surprised at the defence of "prompt engineering" here. It's something easy to do with no real knowledge, and I'm sure having it dismissed hurts some people. But I 100% agree with the author, "prompt engineering" is not science, and I'd say it's not engineering either. All you're doing is exploring the parameter space of particular model in a very crude way. There is no "engineering" going on in this process,…
> But I 100% agree with the author, "prompt engineering" is not science, and I'd say it's not engineering either. All you're doing is exploring the parameter space of particular model in a very crude way. There is no "engineering" going on in this process, just a bunch of trial and error. I wonder what your definition of “science” or “engineering” is…
But prompt engineering is still a pressure point for some people, despite being wildly more simple and accessible (literally tell the thing to do a thing, and if it doesn't do the thing right, reword)
It feels as though we're getting to the technological equivalent of "what IS art anyways", and questions like if non traditional forms like video games are art (I'm thinking all the way up the chain to even say, Madden games)
And in my experience, when something is under constant questioning of whether or not it even counts as X, Y or Z, it usually can technically qualify, but...
If people are constantly debating whether or not it's even X, it's probably just not impressing people who don't engage in it, as opposed to "traditional" concepts of engineering and art, and part of the impression made comes from the investment and irreplaceable skillsets, things few, if anyone else at the time could have done.
This is why taping a banana on the wall is definitely technically art, but not many outside the art community that tapes bananas to walls really think much of it. It's so mundane and accessible a feat that it doesn't garner much merit to passerbys. It's art by the loosest technical definition, and is giving a lot of credit for a small amount of effort anyone could've done.
Admittedly "prompt engineering" is definitely less accessible than a roll of duct tape and a banana but I think we used to just call it "writing/communication", but I guess those who feel capable at that, often just do it manually anyways.
Re: Sorry, but a new prompt for GPT-4 is not a paper
#144Earlier quoted context omitted.
> you realize nobody understands WHY or HOW these models work under the hood right? Of course we understand how they work, we built them! There is no mystery in their mechanisms, we know the number of neurons, their connectivity, everything from the weights to the activation functions. This is not a mystery, this is several decades of technical developments. > it's akin to evolution - we understand the process - that…
> Of course we understand how they work, we built them! There is no mystery in their mechanisms, we know the number of neurons, their connectivity, everything from the weights to the activation functions. This is not a mystery, this is several decades of technical developments. The discovery of DNA’s structure was heralded as containing the same explanatory power as you describe here. Turns out, the story was much mo…
Not at all. It's like saying that since we can read hieroglyphics we know all about ancient Egypt. Deciphering DNA is tool to understand biology, it is not that understanding in itself.
> Turns out, the story was much more complicated then, and is much more complicated now.
We are reverse engineering biology. We are building artificial intelligence. There is a fundamental difference and equating them is fundamentally misunderstanding both of them.
> Anyone today who tells you they know why LLMs are capable of programming, and how they do it, is plainly lying to you.
How so? They can do it because we taught them, there is no magic.
> We have built a complex system that we only understand well at a basic “well there are weights and there’s attention, I guess?” layer. Past that we only have speculation right now.
Exactly in the same way that nobody understand in detail how a complex modern SoC works. Again, there is no magic.
Re: Sorry, but a new prompt for GPT-4 is not a paper
#145Earlier quoted context omitted.
Dude if you think φύσις is free of its own philosophical baggage I got a bridge to sell you.
> its own philosophical baggage Could you elaborate on this statement?
You might retort here "ah well, 'nature' is just the word we use when we speak of observable phenomena in the hard sciences, its not muddied by religion like that crock stuff psychology."
And then I would say, "ok, if 'nature' is just observable phenomena, what is the aim or purpose of the hard sciences? If it is all just observing/experimenting on discrete phenomena, there would be nothing we could do or conclude from the rigor of physics."
You laugh at my insanity (well, if you believed in such a thing): "But we do conclude things from physics, because experiments are reproducible, and with their reproducibility we can gain confidence in generalizing the laws of our universe."
And yes! You would be correct here. But now all the sudden you have committed physics to something just as fundamentally "spiritual" as the soul: that the universe is sensible, rational, and "with laws." Which is indeed just speaking the very same mystical "nature" of ancient Greece from which we get phys-.
But this need not be some damning critique of physics itself (like psychology), and rather, can lead to a higher level understanding of all scientific pursuits: that we are everywhere cursed by a fundamental incompleteness, that in order even to enter into scientific pursuit we must shed an absolute skepticism for a qualified one. Because this is the only way we accumulate a network of reinforced hypotheses and conceptions, which do indeed help us navigate the purely phenomenal world we are bound in.
Re: Sorry, but a new prompt for GPT-4 is not a paper
#146Earlier quoted context omitted.
> you realize nobody understands WHY or HOW these models work under the hood right? Of course we understand how they work, we built them! There is no mystery in their mechanisms, we know the number of neurons, their connectivity, everything from the weights to the activation functions. This is not a mystery, this is several decades of technical developments. > it's akin to evolution - we understand the process - that…
Why stop at chemistry? Chemistry is fundamentally quantum electrodynamics applied to huge ensembles of particles. QED is very well understood and gives the best predictions we have to date of any scientific theory. How come we don’t entirely understand biology then?
Chemistry is indeed applied QED ;) (and you don't need massive numbers of particles to have very complex chemistry)
> How come we don’t entirely understand biology then?
We understand some of the basics (even QED is not reality). That understanding comes from bottom-up studies of biochemistry, but most of it comes from top-down observation of whatever there happens to be around us. The trouble is that we are using this imperfect understanding of the basics to reverse engineer an insanely complex system that involves phenomena spanning 9 orders of magnitude both in space and time.
LLMs did not spawn on their own. There is a continuous progression from the perceptron to GPT-4, each one building on the previous generation, and every step was purposeful and documented. There is no sudden jump, merely an exponential progression over decades. It's fundamentally very different from anything we can see in nature, where nothing was designed and everything appears from fundamental phenomena we don't understand.
As I said, imagining that the current state of AI is anything like biology is a profound misunderstanding of the complexity of both. We like to think we're gods, but we're really children in a sand box.
Re: Sorry, but a new prompt for GPT-4 is not a paper
#147Excuse me? Step by step wasn't paper-worthy? Hard disagree. LLM research is currently in its infancy, because they are no older than a few years old. And a research field in its infancy is bound to have a few noteworthy "no sh*t, Sherlock" papers that would be obvious from hindsight. The fact is, LLMs are a higher-order construct in machine learning, much like a fish is higher-order than a simple cellular colony. Low…
I'm not surprised at the defence of "prompt engineering" here. It's something easy to do with no real knowledge, and I'm sure having it dismissed hurts some people. But I 100% agree with the author, "prompt engineering" is not science, and I'd say it's not engineering either. All you're doing is exploring the parameter space of particular model in a very crude way. There is no "engineering" going on in this process,…
Yes. I come to think of prompt engineering as, in a sense, doing an approximate SELECT query on the latent behavioural space (excuse my lack of proper terminology, my background in ML is pretty thin) that can be thought of as "fishing out" the agent/personality/simulator that is most likely to give you the kind of answer you want. Of course a prompt is a very crude way to explore this space, but to me this is a consequence of extremely poor tooling. For one, llama.cpp now has negative prompts, while the GPT-4 API will probably never have them. So we make-do with the interface available.
> There is no "engineering" going on in this process, just a bunch of trial and error. Perhaps it should be called "prompt guessing."
That is incorrect. It is true that there is a lot of trial and error, yes. But it's not true that it's pure guessing either. While my approach can be best described as a systematic variant of vibe-driven development, at its core it's quite similar to genetic programming. The prompt is mutable, and it's efficacy is possible to evaluate at least in a qualitative sense vs the last version of the prompt. By iterative mutation (rephrasing, restructuring/refactoring the whole prompt, changing out synonyms, adding or removing formatting, adding or removing instructions and contextual information), it is possible to iterate from a terrible initial prompt to a much more elaborate prompt that gets you 90-97% of the way towards nearly exactly what you want to do, by combining the addition of new techniques with subjective judgement on how to proceed (which is incidentally not too different from some strains of classical programming). On GPT-4, at least.
> None of the results of this process will transfer to any other model.
Is that so? Yes, models are somewhat idiosyncratic, and you cannot just drag and drop the same prompt between them. But, in my admittedly limited experience of cross-model prompt engineering, I have found that techniques which helped me to achieve better results with the untuned GPT-3 base model, also helped me greatly with the 7B Llama 1 models. I hypothesise that (in the absence of muddling factors like RLHF-induced censorship of model output), similarly sized models should perform similarly on similar (not necessarily identical) queries. For the time being, this hypothesis is impossible to test because the only realistic peer to GPT-4 (i. e. Claude) is lobotomised to the extent where I would outright pay a premium to not have to use it. I have more to say on this, but won't unless you ask in the interests of brevity.
> Language models have a long history, and while the very large models are impressive, most of their failings have been known for a very long time already. Things like "prompt engineering" will eventually end up in the same graveyard as "keyword engineers" of the past.
Language models have a long history, but a Markov chain can hardly be asked to create a basic Python client for a novel online API. I will also dispute the assertion that we know the "failings" of large language models. Several times now, previously "impossible" tasks have been proven eminently possible by further research and/or re-testing on improved models (better-trained, larger, novel fine-tuning techniques, etc). I am far from being on the LLM hype train, or saying they can do everything that optimists hope they can do. All I'm saying, is that the academia is doing itself a disservice by not looking at the field as something to be explored with no preconceptions, positive or negative.
Re: Sorry, but a new prompt for GPT-4 is not a paper
#148Earlier quoted context omitted.
I'm not surprised at the defence of "prompt engineering" here. It's something easy to do with no real knowledge, and I'm sure having it dismissed hurts some people. But I 100% agree with the author, "prompt engineering" is not science, and I'd say it's not engineering either. All you're doing is exploring the parameter space of particular model in a very crude way. There is no "engineering" going on in this process,…
> But I 100% agree with the author, "prompt engineering" is not science, and I'd say it's not engineering either. All you're doing is exploring the parameter space of particular model in a very crude way. There is no "engineering" going on in this process, just a bunch of trial and error. I wonder what your definition of “science” or “engineering” is…
Re: Sorry, but a new prompt for GPT-4 is not a paper
#149Earlier quoted context omitted.
I’m not sure the scientific method itself can withstand this sort of scrutiny. After all, it’s just making guesses about what will happen and then seeing what happens!
Except there's also, you know, building coherent theories and using those theories to predict the system behavior.
One way to "engage" these internal representations is to include keywords or patterns of text that make that latent knowledge more likely to "activate". Say, if you want to ask about palm trees, include a paragraph talking about a species of a palm tree (no matter whether it contains any information pertaining to the actual query, so long it's "thematically" right) to make a higher quality completion more likely.
It might not be the actual truth or what's going on inside the model. But it works quite consistently when applied to prompt engineering, and produces visibly improved results.
Re: Sorry, but a new prompt for GPT-4 is not a paper
#150Should this be a paper? https://not-just-memorization.github.io/extracting-training-... There is supporting analysis and measurement, but the essence is a single type of prompt, and DeepMind is a heavyweight lab I think it’s fair to say. Moreover there’s evidence people independently reported this result months beforehand on Reddit based on casual observation.
You will notice that they say in the blog post: "If you’re a researcher, consider pausing reading here, and instead please read our full paper for interesting science beyond just this one headline result. In particular, we do a bunch of work on open-source and semi-closed-source models in order to better understand the rate of extractable memorization (see below) across a large set of models." So they are trying to r…
https://arxiv.org/abs/2311.17035
I think it’s valid work, but the original tweet seems to call a prompt based paper into question.
At least enough to clarify where he would stand on an example like this.