Earlier quoted context omitted.
This is about saving money by using the right tool for the job. If you have a system that does a lot of mundane, repetitive work, you don't need a frontier model to do it. It doesn't mean frontier models aren't good at harder tasks.
That's fair and I agree with this framing.
A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
61–70 of 141 posts
Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#62The point that the major labs don’t seem to get is that the vast majority of use cases simply don’t need models that have 50 PhDs and can speak 12 languages. Most use cases are defined within constraints where costs matter a lot. As open weight models and cheap fine tuning services become the norm the whole economic framework of these mega models the labs are in an arms race building just completely crumbles. As does…
what are those cheap fine-tuning services these days?
Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#63What I really started to notice is that SOTA models are really good at putting themselves out of the job. We can see this already with GPT how luna can do 90% of what sol is used for. The only reason why china still bothers 'distilling' models is accurate training data generation, something that oai and anthropic had to spend years collecting while trying to dodge legal challenges. The more intelligent models get, th…
This is a general theme with technology and the 'S-curve'. Let's not even get into whether the improvement for AI reasoning ability has slowed - for practical purposes of writing a React frontend, it has. But other tech is like that - I don't even remember when I bought my LCD TV - 2018 I think? I have no inclination of buying a new one. Technology has a tendency to replace new technology, or intrude into vacant area…
LCD backlights are usually rated for 5-10 years of normal use so your inclination might change soon.
Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#64The point that the major labs don’t seem to get is that the vast majority of use cases simply don’t need models that have 50 PhDs and can speak 12 languages. Most use cases are defined within constraints where costs matter a lot. As open weight models and cheap fine tuning services become the norm the whole economic framework of these mega models the labs are in an arms race building just completely crumbles. As does…
Fine timing takes time and data, though. If smart enough models get cheap enough, then most people have lots of use cases that are cost insensitive enough that it's not worth the effort. Of course, "smart enough" is a low enough threshold for most uses that this is still a problem for the frontier labs. But at the same time, a truly smart enough closed model could also potentially command almost whatever they'd care…
This cycle happens when you have a well defined use case with high volume. It is the far opposite end of the spectrum from the general purpose intelligence on tap that frontier AI models purport to deliver.
I think focused fine tunes and big general models will coexist. Ideally with smart routing and caching to use small, specialized and local options when appropriate.
Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#65The result was: 90% parity achieved with a weird combination of a fine-tuned adapter + 1 deterministic step.
I wrote the whole thing down here: https://alexisrondeau.me/tada/research/FMDiscovery/dashboard... which includes the question, the answer, the 96 experiments and their lineage etc. etc.
Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#66Amazing! I've been in a similar autoresearch-y rabbit-hole lately with getting Apple's 3B Foundation Model to match Sonnet 4.6 on a very specific task. The result was: 90% parity achieved with a weird combination of a fine-tuned adapter + 1 deterministic step. I wrote the whole thing down here: https://alexisrondeau.me/tada/research/FMDiscovery/dashboard... which includes the question, the answer, the 96 experiments…
I like the idea, but it looks really hard to read. Try reading it top to bottom without skipping.
Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#67turns out the bottleneck was never model size, it was having someone who actually understood the problem define the reward function
And I wonder: We've got all these amazing (programming) languages to define solutions; where are the languages to clearly define the problems?
Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#68Amazing! I've been in a similar autoresearch-y rabbit-hole lately with getting Apple's 3B Foundation Model to match Sonnet 4.6 on a very specific task. The result was: 90% parity achieved with a weird combination of a fine-tuned adapter + 1 deterministic step. I wrote the whole thing down here: https://alexisrondeau.me/tada/research/FMDiscovery/dashboard... which includes the question, the answer, the 96 experiments…
> I wrote the whole thing down here I like the idea, but it looks really hard to read. Try reading it top to bottom without skipping.
Is it a layout issue for you (it lays out better on desktop) or the language of the text or...? Let me know!
Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#69Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
#70The point that the major labs don’t seem to get is that the vast majority of use cases simply don’t need models that have 50 PhDs and can speak 12 languages. Most use cases are defined within constraints where costs matter a lot. As open weight models and cheap fine tuning services become the norm the whole economic framework of these mega models the labs are in an arms race building just completely crumbles. As does…
the major labs want to advance science. current business use cases are a happy accident