Earlier quoted context omitted.
The fact that by hand is emphasized so often and so often (see src/handgpt.h comments as well above backward()) makes me think it's AI. Anyways, it looks like gradients derived by hand means they didn't use autograd. They have written out the expression for dL/dW themselves.
I think the readme is clearly LLM generated. The tone, the short sentences, the em dashes, lists, sections.. It all feels like LLM.
An SLM trained on $8 ESP32-S3
21–23 of 23 posts
The author is also replying with an LLM on HN: https://news.ycombinator.com/item?id=49180716
Re: An SLM trained on $8 ESP32-S3
#22Earlier quoted context omitted.
I understand that you are alluring that maximum training-state memory, not parameter count, is the key variable here. So you better start with the smallest model, and go to for highest training data quality, with the outlook of coupling systems together?
You're talking to an LLM, unfortunately https://news.ycombinator.com/threads?id=runtime_lens (enable showdead in your HN settings)
Aw, that is truly annoying...
Thanks for pointing it out!
Until now I never experienced something like this.
Re: An SLM trained on $8 ESP32-S3
#23Would probably be better to demonstrate by exsmple how this approach is used to train on sensor data and then use it (as is hinted by the author) instead of acknowledging that the klingon poc is useless.
I love the cool tech demo! I also would have loved to see a offline sensor calibration package - or whatnot - and analysis of model precision instead of Klingon. Just trying to think up of an actual use case where you would actually truly need an LLM instead of one of the other well known data analytics methods.