Live data from Hacker News

I Squeezed 31.9 TFLOPS out of an RTX 4060, Beating cuBLAS by 14%

github.com

1–2 of 2 posts

Re: I Squeezed 31.9 TFLOPS out of an RTX 4060, Beating cuBLAS by 14%

#2
TL;DR: By instantiating a CUTLASS s16816gemm_f16 kernel with S=3 pipeline stages, 256×128 threadblock tiles, and epilogue vectorization width 4, I achieved 31.9 TFLOPS FP16→FP32 on an RTX 4060 (AD107) at N=8192 — 14.3% faster than the same GPU's cuBLAS baseline of 27.9 TFLOPS, and 52.7% of the 60.55 TFLOPS theoretical tensor-core peak.