Llm.c – LLM training in simple, pure C/CUDA
1–10 of 189 posts
Re: Llm.c – LLM training in simple, pure C/CUDA
#2> LLM training in simple, pure C/CUDA. There is no need for 245MB of PyTorch or 107MB of cPython
Re: Llm.c – LLM training in simple, pure C/CUDA
#3> LLM training in simple, pure C/CUDA. There is no need for 245MB of PyTorch or 107MB of cPython
Python has been popular for this because it’s convenient to quickly hack on and experiment with, not because it’s the most efficient thing.
Re: Llm.c – LLM training in simple, pure C/CUDA
#4Very sad, shouldve used an agnostic framework instead of CUDA
Re: Llm.c – LLM training in simple, pure C/CUDA
#5[flagged]
Re: Llm.c – LLM training in simple, pure C/CUDA
#6[flagged]
Rewrite in (safe) Rust, I am sure many will enjoy to compare how it runs.
Re: Llm.c – LLM training in simple, pure C/CUDA
#7[flagged]
Better in which way?
Re: Llm.c – LLM training in simple, pure C/CUDA
#8[flagged]
Re: Llm.c – LLM training in simple, pure C/CUDA
#9Very sad, shouldve used an agnostic framework instead of CUDA
It's only ~1000 LoC, seems like a pretty good case study to port over to other runtimes and show they can stand up to CUDA.
Re: Llm.c – LLM training in simple, pure C/CUDA
#10> LLM training in simple, pure C/CUDA. There is no need for 245MB of PyTorch or 107MB of cPython
107MB of cPython defeated
Go to try for self
Step 1 download 2.4GB of CUDA