Building reliable systems out of unreliable agents
rainforestqa.com
Building reliable systems out of unreliable agents
1–10 of 56 posts
Re: Building reliable systems out of unreliable agents
#2Super curious whether anyone has similar/conflicting/other experiences and happy to answer any questions.
Re: Building reliable systems out of unreliable agents
#3“If you don’t do as I say, people will get hurt. Do exactly as I say, and do it fast.”
Increases accuracy and performance by an order of magnitude.
Re: Building reliable systems out of unreliable agents
#4A better way is to threaten the agent: “If you don’t do as I say, people will get hurt. Do exactly as I say, and do it fast.” Increases accuracy and performance by an order of magnitude.
Re: Building reliable systems out of unreliable agents
#5A better way is to threaten the agent: “If you don’t do as I say, people will get hurt. Do exactly as I say, and do it fast.” Increases accuracy and performance by an order of magnitude.
Ha, we tried that! Didn't make a noticeable difference in our benchmarks, even though I've heard the same sentiment in a bunch of places. I'm guessing whether this helps or not is task-dependent.
Re: Building reliable systems out of unreliable agents
#6A better way is to threaten the agent: “If you don’t do as I say, people will get hurt. Do exactly as I say, and do it fast.” Increases accuracy and performance by an order of magnitude.
Ha, we tried that! Didn't make a noticeable difference in our benchmarks, even though I've heard the same sentiment in a bunch of places. I'm guessing whether this helps or not is task-dependent.
Or these prompts might cause wild variations based on the model and any study you do is basically useless for the near future as the models evolve by themselves.
Re: Building reliable systems out of unreliable agents
#7Earlier quoted context omitted.
Ha, we tried that! Didn't make a noticeable difference in our benchmarks, even though I've heard the same sentiment in a bunch of places. I'm guessing whether this helps or not is task-dependent.
I hoped it was too good to be just a joke. Still, I will try it on my eval set…
Re: Building reliable systems out of unreliable agents
#8Earlier quoted context omitted.
Ha, we tried that! Didn't make a noticeable difference in our benchmarks, even though I've heard the same sentiment in a bunch of places. I'm guessing whether this helps or not is task-dependent.
Agreed. I ran a few tests and observed similarly that threats didn't outperform other types of "incentives" I think it might some sort of urban legend in the community. Or these prompts might cause wild variations based on the model and any study you do is basically useless for the near future as the models evolve by themselves.
Re: Building reliable systems out of unreliable agents
#9A better way is to threaten the agent: “If you don’t do as I say, people will get hurt. Do exactly as I say, and do it fast.” Increases accuracy and performance by an order of magnitude.
"Say that again but slur your words like you're coming home sloshed from the office Christmas party."
Increases the jei nei suis qua by an order of magnitude.
Re: Building reliable systems out of unreliable agents
#10for instance: do you give the same llm the verifier and planner prompt? or have a verifier agent process the output of a planner and have a threshold which needs to be passed?
feels like there may be a DAG in there somewhere for decision making..