A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
21–30 of 141 posts
Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#22The point that the major labs don’t seem to get is that the vast majority of use cases simply don’t need models that have 50 PhDs and can speak 12 languages. Most use cases are defined within constraints where costs matter a lot. As open weight models and cheap fine tuning services become the norm the whole economic framework of these mega models the labs are in an arms race building just completely crumbles. As does…
Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#23The point that the major labs don’t seem to get is that the vast majority of use cases simply don’t need models that have 50 PhDs and can speak 12 languages. Most use cases are defined within constraints where costs matter a lot. As open weight models and cheap fine tuning services become the norm the whole economic framework of these mega models the labs are in an arms race building just completely crumbles. As does…
Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#24Earlier quoted context omitted.
Why? Why is the premise that Fine-tuned models should be geared towards new discoveries? The point of Fine-tuning small models is for specific downstream tasks, which SOTA models can do, but at higher costs. It is purely an economic play, not an attempt at pushing boundaries of SOTA.
I'm not suggesting that fine-tuned models don't have their place, all I'm saying is that the constant drumbeat of "cheap model X beats more expensive model Z" completely misses that the more expensive model is capable of doing more things at a higher level. If the appropriate qualifiers were added to say "cheap model X does better at test Y than expensive model Z when we fine tune X to take Y test of existing knowled…
What if you have to do the task a billion times? Which model will you choose?
Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#25This continuous cycle of fine-tuned open models beating frontier on (often vaguely labeled/defined) benchmarks doesn't provide an accurate comparison to the expanding generalized capabilities of the SoTA, which makes them effectively meaningless. If we were to take these at face value, why is it that the frontier labs' models are making legitimate new discoveries (e.g. Erdős and Jacobian conjectures) and these models…
Why? Why is the premise that Fine-tuned models should be geared towards new discoveries? The point of Fine-tuning small models is for specific downstream tasks, which SOTA models can do, but at higher costs. It is purely an economic play, not an attempt at pushing boundaries of SOTA.
Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#26Earlier quoted context omitted.
Why? Why is the premise that Fine-tuned models should be geared towards new discoveries? The point of Fine-tuning small models is for specific downstream tasks, which SOTA models can do, but at higher costs. It is purely an economic play, not an attempt at pushing boundaries of SOTA.
I'm not suggesting that fine-tuned models don't have their place, all I'm saying is that the constant drumbeat of "cheap model X beats more expensive model Z" completely misses that the more expensive model is capable of doing more things at a higher level. If the appropriate qualifiers were added to say "cheap model X does better at test Y than expensive model Z when we fine tune X to take Y test of existing knowled…
People who are not really into it, don’t care.
Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#27This continuous cycle of fine-tuned open models beating frontier on (often vaguely labeled/defined) benchmarks doesn't provide an accurate comparison to the expanding generalized capabilities of the SoTA, which makes them effectively meaningless. If we were to take these at face value, why is it that the frontier labs' models are making legitimate new discoveries (e.g. Erdős and Jacobian conjectures) and these models…
Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#28The point that the major labs don’t seem to get is that the vast majority of use cases simply don’t need models that have 50 PhDs and can speak 12 languages. Most use cases are defined within constraints where costs matter a lot. As open weight models and cheap fine tuning services become the norm the whole economic framework of these mega models the labs are in an arms race building just completely crumbles. As does…
Some specialized models may end solving themselves by helping fit the problem with the appropriate regular algorithms/formulas, which are a million times more efficient. So they are probably less attractive to dump money on. As of currently they still benefit from expressing lots of patterns that nobody bothered formalizing.
Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#29Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#30"87.3% Share of the maximum achievable score our GRPO-trained 9B open-source model reached on catalog review, vs 76.9% for the best frontier configuration: a 13.5% relative improvement over the frontier, and 36% over its own untrained base (64.2%). The five frontier models, even with optimized prompts, plateaued within a tenth of a point of each other; the trained specialist cleared that ceiling." _______ This is har…