I find these optimizations fascinating. Anyone familiar with Lemire likely knows about this, but I listened to a podcast episode[1] with him a few days ago and learned about `simdjson`, the tool he authored that parses JSON at 25x the speed of the standard C++ library.[2][3] It's worth looking at if you're into this sort of thing. 1. https://corecursive.com/frontiers-of-performance-with-daniel... 2. https://www.youtu…
yep, great fan of his work here. thanks for sharing that podcast. that kind of optimization requires you to know your machine architecture quite well. SIMD optimizations aren't new. but it's always amazing to see these performance increases on a single machine! our CPUs and GPUs are quite amazing. we have decided, as a field, that we can get enough virtual CPUs, GPUs, or RAM on-demand. and that we shouldn't concern o…
Ontop of this you have to identify what can and cannot be vectorized and how it can be integrated.
Working in simd isn't too hard In itself once you get down to the assembly and toss out all the extra stuff compilers add. If you look at how ffmpeg does it, they just assemble the hand written assembly as the C function itself. Arm64 is very nice to do this in because it has an elegant way in defining vectors and instructions for them.