Live data from Hacker News

Show HN: Tiny-vLLM – high performance LLM inference engine in C++ and CUDA

github.com

1–10 of 26 posts

Re: Show HN: Tiny-vLLM – high performance LLM inference engine in C++ and CUDA

#10
post #2

README is in my opinion (author here) the most interesting - I wrote it to help others build useful mental model to be able to recreate the project yourself, without need to even read my code

Really practical teaching approach. I clicked in to see how safetensors are loaded and just kept reading. Thanks for sharing.
Post reply on HN