Show HN: Fine-tune an 8B model on a 4 GB laptop GPU
31–40 of 47 posts
Re: Show HN: Fine-tune an 8B model on a 4 GB laptop GPU
#32Small open weight local models are the future. While hosted mega models make headlines for doing cool stuff, the vast majority of applications for AI simply don't need all that power, and thus cost. That’s a big part of why businesses are screaming that there’s no ROI from AI. Brining this tech down into small local models is likely where this all converges for the vast majority of use cases and what solves the prese…
If you are into small local models I highly recommend vibe thinker. It's a model trained specifically for reasoning. Basically a problem solver. When compared with other models, on math problems benchmarks, it's closer to models hundred times its size than ten times its size which it beats comfortably. It supports long contexts on limited VRAM and is blazing fast. https://github.com/WeiboAI/VibeThinker
Re: Show HN: Fine-tune an 8B model on a 4 GB laptop GPU
#33Re: Show HN: Fine-tune an 8B model on a 4 GB laptop GPU
#34Earlier quoted context omitted.
Can you just get over it and read what's already there? Anyway, your loss.
If their native language isn’t English then I’ll come off my pedestal
Re: Show HN: Fine-tune an 8B model on a 4 GB laptop GPU
#35They're all very short though. Anyone got a good rule for how much data of this nature is needed to successfully fine-tune a model of this size?
Re: Show HN: Fine-tune an 8B model on a 4 GB laptop GPU
#36Re: Show HN: Fine-tune an 8B model on a 4 GB laptop GPU
#37There are some samples of training data in this folder - https://github.com/MakazhanAlpamys/Soup/tree/main/examples/d... They're all very short though. Anyone got a good rule for how much data of this nature is needed to successfully fine-tune a model of this size?
Re: Show HN: Fine-tune an 8B model on a 4 GB laptop GPU
#38There are some samples of training data in this folder - https://github.com/MakazhanAlpamys/Soup/tree/main/examples/d... They're all very short though. Anyone got a good rule for how much data of this nature is needed to successfully fine-tune a model of this size?
Re: Show HN: Fine-tune an 8B model on a 4 GB laptop GPU
#39This looks cool - what 4 GB GPU laptop do you recommend?
Do not buy a 4 GB card for this. Mine is an RTX 3050 Laptop, I picked it because it is boring hardware that many people already have. If you are buying, buy VRAM. At 0.5B where I could measure both, resident training was 1.43x faster than streaming. And if you do stream, system RAM matters more, the base sits there and has to page-lock. 16 GB is about the floor for 8B.
Re: Show HN: Fine-tune an 8B model on a 4 GB laptop GPU
#40Earlier quoted context omitted.
Can you just get over it and read what's already there? Anyway, your loss.
If their native language isn’t English then I’ll come off my pedestal