Live data from Hacker News

A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

fermisense.com

21–30 of 141 posts

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#22
post #13

The point that the major labs don’t seem to get is that the vast majority of use cases simply don’t need models that have 50 PhDs and can speak 12 languages. Most use cases are defined within constraints where costs matter a lot. As open weight models and cheap fine tuning services become the norm the whole economic framework of these mega models the labs are in an arms race building just completely crumbles. As does…

the major labs want to advance science. current business use cases are a happy accident

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#23
post #13

The point that the major labs don’t seem to get is that the vast majority of use cases simply don’t need models that have 50 PhDs and can speak 12 languages. Most use cases are defined within constraints where costs matter a lot. As open weight models and cheap fine tuning services become the norm the whole economic framework of these mega models the labs are in an arms race building just completely crumbles. As does…

Some specialized models may end solving themselves by helping fit the problem with the appropriate regular algorithms/formulas, which are a million times more efficient. So they are probably less attractive to dump money on. As of currently they still benefit from expressing lots of patterns that nobody bothered formalizing.

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#24

Earlier quoted context omitted.

Why? Why is the premise that Fine-tuned models should be geared towards new discoveries? The point of Fine-tuning small models is for specific downstream tasks, which SOTA models can do, but at higher costs. It is purely an economic play, not an attempt at pushing boundaries of SOTA.

I'm not suggesting that fine-tuned models don't have their place, all I'm saying is that the constant drumbeat of "cheap model X beats more expensive model Z" completely misses that the more expensive model is capable of doing more things at a higher level. If the appropriate qualifiers were added to say "cheap model X does better at test Y than expensive model Z when we fine tune X to take Y test of existing knowled…

> I'm not suggesting that fine-tuned models don't have their place, all I'm saying is that the constant drumbeat of "cheap model X beats more expensive model Z" completely misses that the more expensive model is capable of doing more things at a higher level.

What if you have to do the task a billion times? Which model will you choose?

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#25

This continuous cycle of fine-tuned open models beating frontier on (often vaguely labeled/defined) benchmarks doesn't provide an accurate comparison to the expanding generalized capabilities of the SoTA, which makes them effectively meaningless. If we were to take these at face value, why is it that the frontier labs' models are making legitimate new discoveries (e.g. Erdős and Jacobian conjectures) and these models…

Why? Why is the premise that Fine-tuned models should be geared towards new discoveries? The point of Fine-tuning small models is for specific downstream tasks, which SOTA models can do, but at higher costs. It is purely an economic play, not an attempt at pushing boundaries of SOTA.

Speed play also you can get much faster responses with 9b model.

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#26

Earlier quoted context omitted.

Why? Why is the premise that Fine-tuned models should be geared towards new discoveries? The point of Fine-tuning small models is for specific downstream tasks, which SOTA models can do, but at higher costs. It is purely an economic play, not an attempt at pushing boundaries of SOTA.

I'm not suggesting that fine-tuned models don't have their place, all I'm saying is that the constant drumbeat of "cheap model X beats more expensive model Z" completely misses that the more expensive model is capable of doing more things at a higher level. If the appropriate qualifiers were added to say "cheap model X does better at test Y than expensive model Z when we fine tune X to take Y test of existing knowled…

Maybe because people who are target audience don’t need to have it spelled out like that?

People who are not really into it, don’t care.

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#27

This continuous cycle of fine-tuned open models beating frontier on (often vaguely labeled/defined) benchmarks doesn't provide an accurate comparison to the expanding generalized capabilities of the SoTA, which makes them effectively meaningless. If we were to take these at face value, why is it that the frontier labs' models are making legitimate new discoveries (e.g. Erdős and Jacobian conjectures) and these models…

Depends on use case. That email classifier for legal emails: cheaper at scale as a small tuned model. Let alone better for the planet. Frontier model may have done that tuning!

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#28
post #13

The point that the major labs don’t seem to get is that the vast majority of use cases simply don’t need models that have 50 PhDs and can speak 12 languages. Most use cases are defined within constraints where costs matter a lot. As open weight models and cheap fine tuning services become the norm the whole economic framework of these mega models the labs are in an arms race building just completely crumbles. As does…

Some specialized models may end solving themselves by helping fit the problem with the appropriate regular algorithms/formulas, which are a million times more efficient. So they are probably less attractive to dump money on. As of currently they still benefit from expressing lots of patterns that nobody bothered formalizing.

[deleted]

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#30
post #9

"87.3% Share of the maximum achievable score our GRPO-trained 9B open-source model reached on catalog review, vs 76.9% for the best frontier configuration: a 13.5% relative improvement over the frontier, and 36% over its own untrained base (64.2%). The five frontier models, even with optimized prompts, plateaued within a tenth of a point of each other; the trained specialist cleared that ceiling." _______ This is har…

Why? This is a very narrow task, it’d be surprising if the results were different actually; more interesting question would be how an even smaller model performs in the same finetune.
Post reply on HN