Training a 3.8B LLM to 0.384 CORE for $998
hugovergnes.github.io
Training a 3.8B LLM to 0.384 CORE for $998
1–10 of 24 posts
Re: Training a 3.8B LLM to 0.384 CORE for $998
#2Re: Training a 3.8B LLM to 0.384 CORE for $998
#3Re: Training a 3.8B LLM to 0.384 CORE for $998
#4I don’t blame you for using an LLM to write an article about an LLM that you built. I’ve been reading so much LLM output that now I see it everywhere. I wonder if humans will start writing more like LLMs?
The world of tiny LLMs is so interesting. It is unlocking novel ways to encode information. Why focus on the style of writing instead of the subject matter?
Re: Training a 3.8B LLM to 0.384 CORE for $998
#5I don’t blame you for using an LLM to write an article about an LLM that you built. I’ve been reading so much LLM output that now I see it everywhere. I wonder if humans will start writing more like LLMs?
When we hyper focus on finding something, we find it all the time. The world of tiny LLMs is so interesting. It is unlocking novel ways to encode information. Why focus on the style of writing instead of the subject matter?
Re: Training a 3.8B LLM to 0.384 CORE for $998
#6LLMs are interesting in their own ways but as an engineer, this is a way to unlock a new way of building software.
I recently build a Claude-assisted Excel/CSV parser for a US based property management system (tax compliance). Uses Haiku and has a lot of deterministic code to extract column/row combinations to check known formats and finally handing out the headers to Haiku to give us a translation plan to our support columns.
These would eventually become part of the software, in a tiny LLM. The gap between training (such tiny LLMs) and inference will shrink. We can consult Claude for edge cases, create sample dataset and train a the tiny LLM on demand so we go to Claude less.
The tooling that a project needs is really important. Something I have been feeling as well. Not just in LLM building projects, but regular software projects that are LLM generated.
Re: Training a 3.8B LLM to 0.384 CORE for $998
#7I don’t blame you for using an LLM to write an article about an LLM that you built. I’ve been reading so much LLM output that now I see it everywhere. I wonder if humans will start writing more like LLMs?
Re: Training a 3.8B LLM to 0.384 CORE for $998
#8I don’t blame you for using an LLM to write an article about an LLM that you built. I’ve been reading so much LLM output that now I see it everywhere. I wonder if humans will start writing more like LLMs?
Anyone can level a charge of AI writing, but it's getting to the point that you must write sloppy to seem human. and never never write well with clear technical prose. It is getting to be just ridiculous McCarthyism.
Re: Training a 3.8B LLM to 0.384 CORE for $998
#9I don’t blame you for using an LLM to write an article about an LLM that you built. I’ve been reading so much LLM output that now I see it everywhere. I wonder if humans will start writing more like LLMs?