Live data from Hacker News

A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

fermisense.com

61–70 of 141 posts

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#61

Earlier quoted context omitted.

This is about saving money by using the right tool for the job. If you have a system that does a lot of mundane, repetitive work, you don't need a frontier model to do it. It doesn't mean frontier models aren't good at harder tasks.

That's fair and I agree with this framing.

That was always the framing. I don’t understand your pedantry.

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#62
post #13

The point that the major labs don’t seem to get is that the vast majority of use cases simply don’t need models that have 50 PhDs and can speak 12 languages. Most use cases are defined within constraints where costs matter a lot. As open weight models and cheap fine tuning services become the norm the whole economic framework of these mega models the labs are in an arms race building just completely crumbles. As does…

what are those cheap fine-tuning services these days?

I would not say cheap per say, but in the context of "where is the business success going to be" that's one of the reasons why Mistral focuses on providing tuned model on premises to their customers.

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#63

What I really started to notice is that SOTA models are really good at putting themselves out of the job. We can see this already with GPT how luna can do 90% of what sol is used for. The only reason why china still bothers 'distilling' models is accurate training data generation, something that oai and anthropic had to spend years collecting while trying to dodge legal challenges. The more intelligent models get, th…

This is a general theme with technology and the 'S-curve'. Let's not even get into whether the improvement for AI reasoning ability has slowed - for practical purposes of writing a React frontend, it has. But other tech is like that - I don't even remember when I bought my LCD TV - 2018 I think? I have no inclination of buying a new one. Technology has a tendency to replace new technology, or intrude into vacant area…

> I don't even remember when I bought my LCD TV - 2018 I think? I have no inclination of buying a new one.

LCD backlights are usually rated for 5-10 years of normal use so your inclination might change soon.

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#64
post #48
post #13

The point that the major labs don’t seem to get is that the vast majority of use cases simply don’t need models that have 50 PhDs and can speak 12 languages. Most use cases are defined within constraints where costs matter a lot. As open weight models and cheap fine tuning services become the norm the whole economic framework of these mega models the labs are in an arms race building just completely crumbles. As does…

Fine timing takes time and data, though. If smart enough models get cheap enough, then most people have lots of use cases that are cost insensitive enough that it's not worth the effort. Of course, "smart enough" is a low enough threshold for most uses that this is still a problem for the frontier labs. But at the same time, a truly smart enough closed model could also potentially command almost whatever they'd care…

Yes. Fine tuning is an optimization. Like all optimizations, it’s downstream of figuring out your problem, implementing a naive solution, scaling the naive solution, and getting frustrated by SWAP-C. THEN you start poking at what optimizations are possible, choosing an approach, implementing the optimization, and redeploying.

This cycle happens when you have a well defined use case with high volume. It is the far opposite end of the spectrum from the general purpose intelligence on tap that frontier AI models purport to deliver.

I think focused fine tunes and big general models will coexist. Ideally with smart routing and caching to use small, specialized and local options when appropriate.

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#65
Amazing! I've been in a similar autoresearch-y rabbit-hole lately with getting Apple's 3B Foundation Model to match Sonnet 4.6 on a very specific task.

The result was: 90% parity achieved with a weird combination of a fine-tuned adapter + 1 deterministic step.

I wrote the whole thing down here: https://alexisrondeau.me/tada/research/FMDiscovery/dashboard... which includes the question, the answer, the 96 experiments and their lineage etc. etc.

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#66

Amazing! I've been in a similar autoresearch-y rabbit-hole lately with getting Apple's 3B Foundation Model to match Sonnet 4.6 on a very specific task. The result was: 90% parity achieved with a weird combination of a fine-tuned adapter + 1 deterministic step. I wrote the whole thing down here: https://alexisrondeau.me/tada/research/FMDiscovery/dashboard... which includes the question, the answer, the 96 experiments…

> I wrote the whole thing down here

I like the idea, but it looks really hard to read. Try reading it top to bottom without skipping.

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#67

turns out the bottleneck was never model size, it was having someone who actually understood the problem define the reward function

I agree very much. I'm getting the same take-away even more now that I'm using autoresearch as my main strategy.

And I wonder: We've got all these amazing (programming) languages to define solutions; where are the languages to clearly define the problems?

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#68

Amazing! I've been in a similar autoresearch-y rabbit-hole lately with getting Apple's 3B Foundation Model to match Sonnet 4.6 on a very specific task. The result was: 90% parity achieved with a weird combination of a fine-tuned adapter + 1 deterministic step. I wrote the whole thing down here: https://alexisrondeau.me/tada/research/FMDiscovery/dashboard... which includes the question, the answer, the 96 experiments…

> I wrote the whole thing down here I like the idea, but it looks really hard to read. Try reading it top to bottom without skipping.

Oh, thank you for noticing.

Is it a layout issue for you (it lays out better on desktop) or the language of the text or...? Let me know!

Re: A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

#70
post #13

The point that the major labs don’t seem to get is that the vast majority of use cases simply don’t need models that have 50 PhDs and can speak 12 languages. Most use cases are defined within constraints where costs matter a lot. As open weight models and cheap fine tuning services become the norm the whole economic framework of these mega models the labs are in an arms race building just completely crumbles. As does…

the major labs want to advance science. current business use cases are a happy accident

That party is over. They’re all on the clock now to show they can make money or the plug will be pulled.
Post reply on HN