Earlier quoted context omitted.
It seems like a `ranking_system_prompt` is used to rank the output of other prompts, which is pretty cool! > Your job is to rank the quality of two outputs generated by different prompts. The prompts are used to generate a response for a given task. You will be provided with the task description, the test prompt, and two generations - one for each system prompt. Rank the generations in order of quality. If Generation…
How did they rank the ranking prompt?
GPT-Prompt-Engineer
151–160 of 166 posts
Re: GPT-Prompt-Engineer
#152"Prompt engineering is kind of like alchemy. There's no clear way to predict what will work best. It's all about experimenting until you find the right prompt." lololoollool
I'm honestly curious why you find this funny? Many fields of study have this fuzzy property - it's easier to name which fields don't .
Re: GPT-Prompt-Engineer
#153> Your job is to rank the quality of two outputs generated by different prompts. The prompts are used to generate a response for a given task.
> You will be provided with the task description, the test prompt, and two generations - one for each system prompt.
> Rank the generations in order of quality. If Generation A is better, respond with 'A'. If Generation B is better, respond with 'B'.
> Remember, to be considered 'better', a generation must not just be good, it must be noticeably superior to the other.
> Also, keep in mind that you are a very harsh critic. Only rank a generation as better if it truly impresses you more than the other.
> Respond with your ranking, and nothing else. Be fair and unbiased in your judgement.
So what factors make the "quality" of one prompt "better" than another?
How "impressive" it is to an LLM? What even impresses an LLM? I thought as an AI language model, it lacks human emotional reactions or whatever.
Quality is subjective. Even accuracy is subjective. What needs testing is alignment-- with your interests. The thing is hardcoded to rate based on what aligns with model hosts' interests, not yours.
Only the "classification version" looks capable of making any kind of assertion:
> 'prompt': 'I had a great day!', 'output': 'true' [sentiment analysis I assume?]
The rest of the test prompts aren't even complete sentences, they're half-thoughts you'd expect to hear Peter Gregory mutter to himself:
> 'prompt': 'Launching a new line of eco-friendly clothing' [ok, and?]
The one for 'Why a vegan diet is beneficial for your health' makes some sense at least, but it's really ambiguous.
I'm just some idiot, but if I were creating this, I'd expect the response to ask for a number of expected keywords or something to measure how close each model comes to what the user actually wants. Like, for me, 'what are operating systems' "must" mention all keywords Linux, Windows, and iOS, and "should" mention any of Unix, Symbian, PalmOS, etc.
All tests should tank the score if it detects fourth-wall-breaking "As an AI language model/I don't feel comfortable" crap anywhere in the response. National Geographic got outed on that one the other day.
Re: GPT-Prompt-Engineer
#154This could work really well if it replaced GPT-X-judged performance ranking with human-in-the-loop ranking of prompts, but that’s not as exciting, I guess.
I think the human in the loop evaluation of prompts is a red herring since the LLMs were trained with HFRL, which optimized them to be convincing. Humans aren't reliable for evaluating these things because it was literally trained to make us think it's doing well.
Re: GPT-Prompt-Engineer
#155Earlier quoted context omitted.
I'm honestly curious why you find this funny? Many fields of study have this fuzzy property - it's easier to name which fields don't .
May I ask, are you a VC? This is a complete non sequitur.
It's a response to the notion that the quote above comparing writing prompts to alchemy is so wrong that it's funny in some way. I thought it was a pretty good analogy.
Re: GPT-Prompt-Engineer
#156Earlier quoted context omitted.
Software engineer has always been a daft, grandiose term that seems aimed at prestige rather than reality in the majority of cases. If coders can call themselves engineers, no reason why anybody else solving puzzles for a living can't.
I’m guessing you aren’t from the US because in the US there is not much prestige in the title of engineer, but it seems to be a title people in Europe get weirdly hung up on. If anything, software engineering positions are more prestigious than most traditional engineering positions here.
Knot engineer, floor polishing engineer, water spillage cleanup engineer, tax minimization engineer, heart engineer, bus operating engineer, flower pruning engineer
Re: GPT-Prompt-Engineer
#157Re: GPT-Prompt-Engineer
#158Isn’t engineering an exact science while prompt engineering is completely not? Although, even software engineering being an exact science, it is a funny one: most of us don’t get certified as like, let’s say, mechanical engineers do. Would they say we are engineers? So perhaps the “engineer” term got overloaded in recent years?
Prompt "engineering" is just writing prayers to forest faeries. Whilst BASIC/JavaScript/etc are all magic incantations to a child, a child will soon figure out there's underlaying logic, and learn the ability to reason about what code does, and what certain changes will do. With prompts, it's all faerie logic. There is nothing to learn, there are only magic incantations that change drastically if the model is updated…
We need to move away from prompt-engineering - it's AI-Management. You pretend you're speaking to another (albeit confusing/confused) person when extracting work from a model. You're coaxing things out of it based on hearsay and mysticism that work most of the time. Sounds a lot like AGILE and free pizza to get a junior to stay late and deliver on time.
That's not engineering, that's management.
Re: GPT-Prompt-Engineer
#159Earlier quoted context omitted.
Does the scientific method include peer review? I thought that was a recent phenomenon.
Peer review has been a thing since, albeit in a much less rigorous way, at least the XVII century. The so-called 'republic of letters' was a world (europe) wide society of philosophers, mathematicians, and early experimenters sharing results and data in writing. Mathematicians have been publishing proofs and challenging each other in Europe since at least the Italian Risorgimento...
Re: GPT-Prompt-Engineer
#160Earlier quoted context omitted.
That's if you're trying to become a scientist. You can do science even if your experiment doesn't contain any of those elements besides trying multiple things and observing the outcomes.
Sure, the same way I can practice medicine by putting leeches on people. I'm not trying to become a physician so it's alright to call that medicine.
https://www.nationalgeographic.com/magazine/article/leeches-...