Live data from Hacker News

I used RL fine-tuning to make an LLM generate ugly and unpythonic FizzBuzz code

seantey.github.io

1–2 of 2 posts

Re: I used RL fine-tuning to make an LLM generate ugly and unpythonic FizzBuzz code

#2
I wrote up a blog post for a hackathon project where I used RL fine-tuning to make an LLM generate intentionally ugly and unpythonic FizzBuzz code. The post covers what I learned about reward shaping and GRPO. Feedback on the writing or content is welcome!