Earlier quoted context omitted.
why...?
I guess we found the target audience for models that slather READMEs in emojis.
Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
21–30 of 54 posts
Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#22No idea why you got downvoted into oblivion with the context post. Cool idea!
Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#23Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#24Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#25I'm curious to whether the recursively trained models degenerate to troglodytes after a couple of generations.
Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#26AI trains AI already, agents are happy to spin up real training pipelines for deep learning or regression models or whatever you want right? I guess the advantage to your project is that it provides a framework to allow the agent to access extra compute?
Yes I'd heard the labs (Anthropic mostly) speaking about LLMs training LLMs, so I wanted to make things a little more concrete and test it out myself! Essentially you are correct though, my framework allows the agent access to compute, but also the agent itself is being trained to become better at training models with that compute.
Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#27Earlier quoted context omitted.
Yes I'd heard the labs (Anthropic mostly) speaking about LLMs training LLMs, so I wanted to make things a little more concrete and test it out myself! Essentially you are correct though, my framework allows the agent access to compute, but also the agent itself is being trained to become better at training models with that compute.
i remeber reading in one of the release blog posts that that version was the "first that codex helped train"
Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#28Earlier quoted context omitted.
I think the counter point for these projects is that you may not need a deep understanding if you can measure the outcome. While this may not be true every time today, it plausibly will be in the future - making the activity worthwhile.
Well, you say that, but when "measuring" anything in RL, that measurement itself is not always obvious. That is, creating the scoring system/judge models etc for RL is not easy at all. You can easily create an RL loop which is getting better and improving its scores, but actually the result is totally garbage, because you're measuring the wrong thing.
Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#29Lots of emoji in that readme. Was it mainly codex?
Mainly Fable, but It was me who wanted to emojis added hah. I also of course edited the README by hand (crazy I know), but the code is entirely fable
Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#30 –$1.3k -> ~$1.3k