I want to make a Discord bot that impersonates all my friends and continues to refine the model as the conversations continue. Basically this [1] post, but with a more modern model and, ideally, reinforcement learning. Seems like this would fit the bill.... Is there anything else that would make this easier? [1] https://www.izzy.co/blogs/robo-boys.html
Show HN: LlamaGym – fine-tune LLM agents with online reinforcement learning
21–29 of 29 posts
Re: Show HN: LlamaGym – fine-tune LLM agents with online reinforcement learning
#22I have a PC that is able to run e.g. Mistral Instruct 7B Q4 inference with around 30 token/s.
How (computation and memory) expensive would it be to also run backpropagation in addition to inference?
I'm aware that the models are typically fed with much more and better data than what is typically provided during normal conversations but on the other hand if I could finetune my local model a teeny tiny bit during during / after each conversation I have with it anyways, it would after a while be perfectly customize for me.
I'm also aware that this could be problematic for models that are used by multiple users but my intended use case would be personal use by a single user.
Re: Show HN: LlamaGym – fine-tune LLM agents with online reinforcement learning
#23From the title I misunderstood what it does. However, now I'm wondering if what I thought is was (don't ask my why I thought it) is possible: I have a PC that is able to run e.g. Mistral Instruct 7B Q4 inference with around 30 token/s. How (computation and memory) expensive would it be to also run backpropagation in addition to inference? I'm aware that the models are typically fed with much more and better data than…
AFAIK the model can’t be quantized during backprop, so right there you’d need a ton of RAM.
Backprop is faster bc it can be parallelized, but IIRC you need to hold an entire copy of the model for each backprop process.
Re: Show HN: LlamaGym – fine-tune LLM agents with online reinforcement learning
#24Re: Show HN: LlamaGym – fine-tune LLM agents with online reinforcement learning
#25When 150 lines of boilerplate can land you the first page on HN, maybe it is, in fact, the end of programming?
Carmack's infamous fast inverse square root was only 13 lines. Measuring code by line metrics rather than its contents reflects shallow, questionable comprehension.
Re: Show HN: LlamaGym – fine-tune LLM agents with online reinforcement learning
#26From the title I misunderstood what it does. However, now I'm wondering if what I thought is was (don't ask my why I thought it) is possible: I have a PC that is able to run e.g. Mistral Instruct 7B Q4 inference with around 30 token/s. How (computation and memory) expensive would it be to also run backpropagation in addition to inference? I'm aware that the models are typically fed with much more and better data than…
Very expensive. AFAIK the model can’t be quantized during backprop, so right there you’d need a ton of RAM. Backprop is faster bc it can be parallelized, but IIRC you need to hold an entire copy of the model for each backprop process.
Re: Show HN: LlamaGym – fine-tune LLM agents with online reinforcement learning
#27Re: Show HN: LlamaGym – fine-tune LLM agents with online reinforcement learning
#28From the title I misunderstood what it does. However, now I'm wondering if what I thought is was (don't ask my why I thought it) is possible: I have a PC that is able to run e.g. Mistral Instruct 7B Q4 inference with around 30 token/s. How (computation and memory) expensive would it be to also run backpropagation in addition to inference? I'm aware that the models are typically fed with much more and better data than…