Viewing profile — simon_vtr
simon_vtr
HN member- Joined
- Sat, Jun 07, 2025, 1:42 PM UTC
- HN karma
- 6
- Public activity
- 4 items
- HN profile
- View on Hacker News ↗
About simon_vtr
No profile information was provided.
Recent public activity
-
comment
Comment #44397160
yea, i wrote this blogpost rather to show how to use scan in different ways than the canonical example of calculating prefix sum of a vector shown in introductions on gpu programmi…
-
comment
Comment #44210545
It doesn’t. The batch size is just 8. This is a very good trick and often needed to archive peak performance in memory bound kernels. You can checkout the equivalent code in cuda a…
-
comment
Comment #44209660
That was exactly my reason to write this blogpost and optimise transpose. It is a simple educational yet not trivial example to learn the basics.
-
comment
Comment #44209636
The kernels I mention in CUDA use all the equivalent logic like the Mojo kernels. You can find them on my GitHub: https://github.com/simveit/effective_transpose You may want to pro…