Live data from Hacker News

My finetuned models beat OpenAI's GPT-4

mlops.systems

81–90 of 98 posts

Re: My finetuned models beat OpenAI's GPT-4

#81

(Disclaimer: I'm the founder of OpenPipe, one of the fine-tuning services OP tried and ultimately the one that produced the highest performing model, it appears.) Data extraction is a use case that fine-tuned models are fantastic at, so I'm not surprised that OP got good results. That said, I've also found it's pretty easy to beat GPT-4 across many task types if you have a way of getting strong training data. We publ…

[dead]

Re: My finetuned models beat OpenAI's GPT-4

#82
post #68

Earlier quoted context omitted.

The way you cut that quote turns it into an assertion that doesn't exist in parent post. They didn't make the (incorrect) statement that no other serious, useful application exists. But that's how it reads when you cut off before "I have actually engaged in for real work and found useful"

To be fair the original sentence could still be implying the same thing. The second half of the sentence just sounds like a hedge.

Well I precisely talked about things I have engaged professionally. Obviously this cannot cover everything one may do, eg I do not build chatbots for customer service or stuff like that, thus I obviously cannot speak for all possible applications of LLMs and how useful they may be. I am pretty sure there will be useful applications in fields I am not and will not be engaged in as nobody engages with everything. However, some other things that I have tried (eg copilots, summarising scientific articles) imo create much more hype than real value. They can be a bit useful if you know what to actually use them for and what their limits are, but nowhere close to the hype they generate, and I just find myself just googling again tbh. They are absolutely horrible especially with more niche subjects and areas. On the other hand, data extraction and structuring has a quite universal application, has already demonstrated usefulness and potential, and seems a quite realistic, down to earth application that I am happy to see other people and startups working on. Not as fancy, and harder to build hype upon, but very useful regardless.

Re: My finetuned models beat OpenAI's GPT-4

#83

(Disclaimer: I'm the founder of OpenPipe, one of the fine-tuning services OP tried and ultimately the one that produced the highest performing model, it appears.) Data extraction is a use case that fine-tuned models are fantastic at, so I'm not surprised that OP got good results. That said, I've also found it's pretty easy to beat GPT-4 across many task types if you have a way of getting strong training data. We publ…

Is this something, as a tech enthusiast that's no expert, I can easily fine tune are run? My use case would be fine tuning on technical docs. Specific news, 2 years of blog posts, primary source material, and Twitter explainer thread. I want to gather all the niche information of a topic from the last two years, dump it into this and have an LLM that is a subject-matter expert.

Here is an example of the Predibase platform, referred in the article for the Solar model, but that can train also Llama-3, Phi-3 and Mistral. https://www.youtube.com/watch?v=R2JQhzfaOFw&themeRefresh=1 I think you can assess by yourself if it's easy enough to do for you. (Predibase founder here)

Re: My finetuned models beat OpenAI's GPT-4

#84
post #14
post #2

And that’s the point of fine tuning models. Still good to see someone walk through their fine tuning process, with a mix of hosted and local options.

On that note: is there a good service for “here’s my dataset”, please fine tune these 9 models and give me evaluation stats?

Predibase ( http://predibase.com ), also referred in the article, is a platform specifically designed for exactly that. It also has "repos" for finetuning multiple models and comapre their performance and keeping things organzie. It also allow you to query any of the finetuned models on the fly from a single GPU with multi-lora serving. (Predibase founder here)

Re: My finetuned models beat OpenAI's GPT-4

#85
We got very similar findings: we published a paper that show that smaller LLMs (3-7b) when finetuned with LoRA can match or outperform GPT-4 on a variety of tasks (29 out of 31) including classification, summarization, info extraction, "reasoning". https://arxiv.org/abs/2405.00732 (Predibase cofounder and coauthor of the paper)

Re: My finetuned models beat OpenAI's GPT-4

#86
At Predibase, we recently conducted 700+ fine-tuning experiments to benchmark the performance of popular open-source LLMs across 30 tasks and compared their results to GPT-4.

85% of the time they beat GPT-4.

You can see the results here: https://predibase.com/fine-tuning-index.

The site has a series of interactive charts and a link to our Arxiv paper.

Re: My finetuned models beat OpenAI's GPT-4

#87

(Disclaimer: I'm the founder of OpenPipe, one of the fine-tuning services OP tried and ultimately the one that produced the highest performing model, it appears.) Data extraction is a use case that fine-tuned models are fantastic at, so I'm not surprised that OP got good results. That said, I've also found it's pretty easy to beat GPT-4 across many task types if you have a way of getting strong training data. We publ…

Is this something, as a tech enthusiast that's no expert, I can easily fine tune are run? My use case would be fine tuning on technical docs. Specific news, 2 years of blog posts, primary source material, and Twitter explainer thread. I want to gather all the niche information of a topic from the last two years, dump it into this and have an LLM that is a subject-matter expert.

Fine tuning doesn't quite work that way. You have to format the training data set as request/response. The idea of fine tuning is to get the model to output things in a specific format, style or structure.

Your use case is better suited to RAG. This is where you retrieve data from a large dataset and inject it into the user's request so the AI model has the context it needs to answer accurately.

But that's not a silver bullet and you would need to spend significant time on chunking strategy and ranking of results to hopefully get a decent response accuracy.

Re: My finetuned models beat OpenAI's GPT-4

#88
post #80
post #53

Earlier quoted context omitted.

It seems to me this means whoever has hoarded and declared ownership of the most personal data will make the best products. Kinda like how some people liked their targeted ads because they’re more “relevant”, only now it’s not just ads but useful products. Another winner is of course platform owners like Apple and Microsoft who can scrape your data off their apps and products, even locally. This is a much bigger edge…

Your end point I think is exactly right. I think your first one is getting downvoted hard because your first sentence is not at all how any of this works. Sucking down personal data isn't JUST a bad idea for privacy, it's actually also bad for "making the best products," I think you're overstating the extent to which all that data that is stolen and sold to the highest bidder actually helps the company buying it?

Ah thanks for pointing out. I don't care much for LLMs at all, but my point was simply that whoever has data, and especially personalized data, has an upper hand in making LLMs into better end user product, for those that like them. This may be underestimated right now when most dick measuring is comparing model-model not integration into a product.

> data that is stolen and sold to the highest bidder

Didn’t mean necessarily the data brokers (although that’s an interesting angle), but say Apple now has a bunch of info about your calendar, email, contacts, then clearly they have an upper hand in providing better products than an anonymous API call. Not all products need personalization but LLMs? I can think of tons of use cases.

Re: My finetuned models beat OpenAI's GPT-4

#90
Eventually people will realize any underdetermined system of equations has infinitely many solutions. Give me any open source AI model and I will beat any SOTA benchmark. Why am I so confident? Because curve fitting can be applied to any data set to get as good of a result as needed. Combine this approach with mixtures of "experts" and any predetermined set of benchmarks will fall to a curve fit to the benchmark.

The hype is really getting tiresome. There is no way to get from here to any intelligent system with the current techniques. New breakthroughs will require insights into discrete spaces which are not amenable to curve fitting with gradient descent.

Post reply on HN