Live data from Hacker News

Fusing a 27B ternary LLM's whole decode step into one CUDA kernel

twitter.com

1–2 of 2 posts