Live data from Hacker News

OpenAI Reinforcement Fine-Tuning Research Program

openai.com

61–67 of 67 posts

Re: OpenAI Reinforcement Fine-Tuning Research Program

#61
post #28

Earlier quoted context omitted.

They're searching for enterprise customers before they become a commodity.

Llama 3.3 is insanely good and can be run on a Mac mini with 64GB of ram for $2k USD. OpenAI is screwed. (As an aside: very interesting Google tried to go closed source and objectively lost the race, and Meta went open and is the real threat to OpenAI.)

I haven't tried it on my $150 2080ti, but I know someone running it on a 3060 and it's not that horrible. Wild times.

Those M4 Macs with shared RAM definitely seem to be the best way to go for this, though.

Re: OpenAI Reinforcement Fine-Tuning Research Program

#62
post #58

Earlier quoted context omitted.

With 64GB you only get a lower quality quantized version.

That’s the one I’m using. So far it’s quite good, and when I gave it and Claude the same programming problem not only did Llama give a better result, when I showed that result to Claude it also said the Llama approach was better. Claude is already better than GPT on average at coding, so yeah, bad news for OpenAI as Llama is now potentially better at coding. Of course Meta has a propriety training set of extremely hi…

That's great to hear. I just want to make sure that you're aware that you're not getting the 100% FP16 experience. I guess at 8bit it's still pretty much the same.

Re: OpenAI Reinforcement Fine-Tuning Research Program

#63

Earlier quoted context omitted.

You are using the terms "uncensored" "malicious" and "unaligned" interchangeably. There would appear to be a few issues with that, the most obvious being the uncensored model would presumably be "aligned" with what the finetuner wants.

I didn't use two of those three terms, so maybe confirming you read the comment you replied to is in order? "Uncensored" is a broad phrase but those in post-training community who post-train "uncensored" versions of a models have a very specific meaning: the creator is stripping refusals. They do it via techniques like abliteration, or SFT on "toxic" datasets, but the toxic datasets tend to be low quality answers and…

>>>I didn't use two of those three terms, so maybe confirming you read the comment you replied to is in order?

You replied to a question asking someone to elaborate on "malicious fine tuning". Specifically someone asked for elaboration on "Dawn Song was very clear that malicious fine tuning is a top priority among implementers right now."

Whatever your actual intent, it's only natural that I read your comment on "uncensored" models as an explanation of "malicious fine tuning".

The parent comment about "malicious fine tuning" remains unexplained. Since nobody else replied, I suppose we will never know how this Dawn Song person defines "malicious".

Re: OpenAI Reinforcement Fine-Tuning Research Program

#64

Earlier quoted context omitted.

I didn't use two of those three terms, so maybe confirming you read the comment you replied to is in order? "Uncensored" is a broad phrase but those in post-training community who post-train "uncensored" versions of a models have a very specific meaning: the creator is stripping refusals. They do it via techniques like abliteration, or SFT on "toxic" datasets, but the toxic datasets tend to be low quality answers and…

>>>I didn't use two of those three terms, so maybe confirming you read the comment you replied to is in order? You replied to a question asking someone to elaborate on "malicious fine tuning". Specifically someone asked for elaboration on "Dawn Song was very clear that malicious fine tuning is a top priority among implementers right now." Whatever your actual intent, it's only natural that I read your comment on "unc…

A ton of words to not just admit you lost track of who you were replying to.

Re: OpenAI Reinforcement Fine-Tuning Research Program

#65

Earlier quoted context omitted.

>>>I didn't use two of those three terms, so maybe confirming you read the comment you replied to is in order? You replied to a question asking someone to elaborate on "malicious fine tuning". Specifically someone asked for elaboration on "Dawn Song was very clear that malicious fine tuning is a top priority among implementers right now." Whatever your actual intent, it's only natural that I read your comment on "unc…

A ton of words to not just admit you lost track of who you were replying to.

I did not lose track. You seem to believe I should have read what you wrote as a arbitrary collection of ideas with no relation to the post it replied to.

Re: OpenAI Reinforcement Fine-Tuning Research Program

#66

Earlier quoted context omitted.

A ton of words to not just admit you lost track of who you were replying to.

I did not lose track. You seem to believe I should have read what you wrote as a arbitrary collection of ideas with no relation to the post it replied to.

If you can't see the relation maybe you need to understand the topic a bit better before diving head into conversations about it...
Post reply on HN