Live data from Hacker News

Building an LLM from Scratch: Automatic Differentiation (2023)

bclarkson-code.github.io

1–10 of 18 posts

Re: Building an LLM from Scratch: Automatic Differentiation (2023)

#2
As a chronic premature optimizer my first reaction was, "Is this even possible in vanilla python???" Obviously it's possible, but can you train an LLM before the heat death of the universe? A perceptron, sure, of course. A deep learning model, plausible if it's not too deep. But a large language model? I.e. the kind of LLM necessary for "from vanilla python to functional coding assistant."

But obviously the author already thought of that. The source repo has a great motto: "It don't go fast but it do be goin'" [1]

I love the idea of the project and I'm curious to see what the endgame runtime will be.

[1] https://github.com/bclarkson-code/Tricycle

Re: Building an LLM from Scratch: Automatic Differentiation (2023)

#4
post #2

As a chronic premature optimizer my first reaction was, "Is this even possible in vanilla python???" Obviously it's possible , but can you train an LLM before the heat death of the universe? A perceptron, sure, of course. A deep learning model, plausible if it's not too deep. But a large language model? I.e. the kind of LLM necessary for "from vanilla python to functional coding assistant." But obviously the author a…

Why wouldn't it be possible? You can generate machine code with Python and call into it with ctypes. All your deep learning code is still in Python, but in the runtime it gets JIT compiled into something faster.

Re: Building an LLM from Scratch: Automatic Differentiation (2023)

#5
Every one should go through this rite of passage work and get to the "Attention is all you need" implementation. It's a world where engineering and the academic papers are very close and reproducible and a must for you to progress in the field.

(see also andre karpathys zero to hero nn series on youtube as well its very good and similar to this work)

Re: Building an LLM from Scratch: Automatic Differentiation (2023)

#6
post #5

Every one should go through this rite of passage work and get to the "Attention is all you need" implementation. It's a world where engineering and the academic papers are very close and reproducible and a must for you to progress in the field. (see also andre karpathys zero to hero nn series on youtube as well its very good and similar to this work)

+1 for Karpathy, the series is really good

Re: Building an LLM from Scratch: Automatic Differentiation (2023)

#8
post #5

Every one should go through this rite of passage work and get to the "Attention is all you need" implementation. It's a world where engineering and the academic papers are very close and reproducible and a must for you to progress in the field. (see also andre karpathys zero to hero nn series on youtube as well its very good and similar to this work)

I would also recommend going through Callum McDougall/Neel Nanda's fantastic Transformer from Scratch tutorial. It takes a different approach to conceptualizing the model (or at least, it implements it in a way which emphasizes different characteristics of Transformers and self-attention), which I found deeply satisfying when I first explored them.

https://arena-ch1-transformers.streamlit.app/%5B1.1%5D_Trans...

Re: Building an LLM from Scratch: Automatic Differentiation (2023)

#9
post #7

is there an existing SLM that resembles an LLM in architecture that includes the code for training it ? i realize the cost and time to train may be prohibitive and that quality on general english might be very limited, but is the code itself available ?

Not sure what you mean with SLM, but https://github.com/karpathy/nanoGPT
Post reply on HN