CUDA-l2: Surpassing cuBLAS performance for matrix multiplication through RL
1–10 of 17 posts
Re: CUDA-l2: Surpassing cuBLAS performance for matrix multiplication through RL
#2Re: CUDA-l2: Surpassing cuBLAS performance for matrix multiplication through RL
#3Re: CUDA-l2: Surpassing cuBLAS performance for matrix multiplication through RL
#4[flagged]
Re: CUDA-l2: Surpassing cuBLAS performance for matrix multiplication through RL
#5Re: CUDA-l2: Surpassing cuBLAS performance for matrix multiplication through RL
#6They claim the algorithm "discovered" the new techniques, but the methods described in section 5 do not seem all that novel to me. It smells like it could be "laundering" the literature [1] and reshuffling existing techniques. This is not inherently a bad thing, but I would hope that if it is borrowing existing techniques, the appropriate citation would eventually make it into this paper. [1]: https://www.argmin.net/…
Re: CUDA-l2: Surpassing cuBLAS performance for matrix multiplication through RL
#7Re: CUDA-l2: Surpassing cuBLAS performance for matrix multiplication through RL
#8They claim the algorithm "discovered" the new techniques, but the methods described in section 5 do not seem all that novel to me. It smells like it could be "laundering" the literature [1] and reshuffling existing techniques. This is not inherently a bad thing, but I would hope that if it is borrowing existing techniques, the appropriate citation would eventually make it into this paper. [1]: https://www.argmin.net/…
Re: CUDA-l2: Surpassing cuBLAS performance for matrix multiplication through RL
#9They claim the algorithm "discovered" the new techniques, but the methods described in section 5 do not seem all that novel to me. It smells like it could be "laundering" the literature [1] and reshuffling existing techniques. This is not inherently a bad thing, but I would hope that if it is borrowing existing techniques, the appropriate citation would eventually make it into this paper. [1]: https://www.argmin.net/…
Re: CUDA-l2: Surpassing cuBLAS performance for matrix multiplication through RL
#10[flagged]
This is a standard which few kernels will ever meet. I'd say requiring a numerical proof is the same as requiring no proof at all - because it won't ever happen unless you're validating silicon or something equally expensive.