Batched reward model inference and Best-of-N sampling #1 Post by rawsh » Tue, Nov 19, 2024, 6:19 AM UTC Batched reward model inference and Best-of-N samplingraw.sh