Why we write our own C and C++ inference engines
1–10 of 63 posts
Re: Why we write our own C and C++ inference engines
#2Re: Why we write our own C and C++ inference engines
#3Re: Why we write our own C and C++ inference engines
#4Should have started with writing your own blog posts.
Re: Why we write our own C and C++ inference engines
#5Wasm size from 30Mb to 300kb and 1.5x speedup. It's definitely worth it for performance or distribution size.
Re: Why we write our own C and C++ inference engines
#6Re: Why we write our own C and C++ inference engines
#7Should have started with writing your own blog posts.
Re: Why we write our own C and C++ inference engines
#8But the cpp port of vllm looks great, that'd be great if you'll maintain that. I hit the same limitations with vllm.
Re: Why we write our own C and C++ inference engines
#9Should have started with writing your own blog posts.
I went through the post because of your comment but it really doesn't look like AI slop. Can you please share why you feel like its slop and not written by a human? I can also say "should have started writing your own comments" to you and its unfalsifiable. Blanket accusations with no proof is not a good move really.
> The method, the measurements, and what it costs us.
> That is the general shape of these wins.
> Parity is the gate, speed is the follow-up
I could go on and on, but you probably get the point. If you don't find anything funny with the above, you might have not been enough-exposed to slop.
Re: Why we write our own C and C++ inference engines
#10Should have started with writing your own blog posts.