AI trains AI already, agents are happy to spin up real training pipelines for deep learning or regression models or whatever you want right? I guess the advantage to your project is that it provides a framework to allow the agent to access extra compute?
Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
31–40 of 54 posts
Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#32Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#33AI trains AI already, agents are happy to spin up real training pipelines for deep learning or regression models or whatever you want right? I guess the advantage to your project is that it provides a framework to allow the agent to access extra compute?
Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#34Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#35I think that En Dash is supposed to be Tilde –$1.3k -> ~$1.3k
Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#36Earlier quoted context omitted.
Mainly Fable, but It was me who wanted to emojis added hah. I also of course edited the README by hand (crazy I know), but the code is entirely fable
i read around launch that anthropic will fallback to opus if fable is used for frontier LLM development. did you run into anything like that?
Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#37Earlier quoted context omitted.
Did you read the README?
The AI generated README?
Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#38Lots of emoji in that readme. Was it mainly codex?
“The Anthropic team for their incredible coding models (Fable-5 wrote every line of code in this project), and the Claude Code harness.” Source: the repo
Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#39AI trains AI already, agents are happy to spin up real training pipelines for deep learning or regression models or whatever you want right? I guess the advantage to your project is that it provides a framework to allow the agent to access extra compute?
Not sure how much of it was marketing but M2.7 supposedly trained itself https://www.minimax.io/blog/minimax-m27
Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
#40AI trains AI already, agents are happy to spin up real training pipelines for deep learning or regression models or whatever you want right? I guess the advantage to your project is that it provides a framework to allow the agent to access extra compute?
I guess more important than LLMs training LLMs (which is mostly the same code over and over) is LLMs cleaning and curating training data.