Live data from Hacker News

Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)

github.com

51–54 of 54 posts

Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)

#51
post #37

Earlier quoted context omitted.

The AI generated README?

Your comments are casting aspersion without showing that you have looked into the work or its author. This project is a case study in good work done with heavy ai assistance while it is clear that there is a skilled person leading. I predict that this will become very common and welcome here.

Why should anyone put serious time and effort into using/understanding a product when the author hasn't put serious time and effort into making it?

I could make this exact same thing over a weekend and post it on Hackernews. But I won't because I would be embarrassed to do so.

The bar for posting something to HN should be high, the bar for wanting people to read your code, your writing, should be putting serious effort and thought into it. Not just vibe coding something up with a vibe coded README and 100% vibe coded code and not even a novel idea or implementation.

Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)

#52
post #37

Earlier quoted context omitted.

Your comments are casting aspersion without showing that you have looked into the work or its author. This project is a case study in good work done with heavy ai assistance while it is clear that there is a skilled person leading. I predict that this will become very common and welcome here.

Why should anyone put serious time and effort into using/understanding a product when the author hasn't put serious time and effort into making it? I could make this exact same thing over a weekend and post it on Hackernews. But I won't because I would be embarrassed to do so. The bar for posting something to HN should be high, the bar for wanting people to read your code, your writing, should be putting serious effo…

I think you are putting more effort into arguing it’s not worth it to read the link, than just reading the link would take. I am not sure what you think your comments are accomplishing, they are not useful, just pedantic

Re: Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)

#53
post #11

Earlier quoted context omitted.

I think the counter point for these projects is that you may not need a deep understanding if you can measure the outcome. While this may not be true every time today, it plausibly will be in the future - making the activity worthwhile.

Well, you say that, but when "measuring" anything in RL, that measurement itself is not always obvious. That is, creating the scoring system/judge models etc for RL is not easy at all. You can easily create an RL loop which is getting better and improving its scores, but actually the result is totally garbage, because you're measuring the wrong thing.

They said that the we are moving towards " `define good for the model` (success criteria/rubrics)" and you've just gone no no, you don't understand you need to measure the right thing. Feels like you've just completely agreed with them?
Post reply on HN