Gemlite: Towards Building Custom Low-Bit Fused CUDA Kernels
mobiusml.github.io
Gemlite: Towards Building Custom Low-Bit Fused CUDA Kernels
1–3 of 3 posts
Re: Gemlite: Towards Building Custom Low-Bit Fused CUDA Kernels
#2This would be great to have for the Triton language as well.
Re: Gemlite: Towards Building Custom Low-Bit Fused CUDA Kernels
#3Weird that they don't mention Triton? I only skimmed it, but I'm not sure what the pros and cons would be vs. Triton, which is the tool I'd use if I wanted custom quantized inference kernels.