I wonder if the 4x performance boost of 128-bit SIMD over 32-bit would drop to 2x if WebAssembly added 64-bit registers/instructions.
[1] https://github.com/awelm/simd-wasm-profiling/blob/master/fil...
11–20 of 25 posts
I wonder if the 4x performance boost of 128-bit SIMD over 32-bit would drop to 2x if WebAssembly added 64-bit registers/instructions.
[1] https://github.com/awelm/simd-wasm-profiling/blob/master/fil...
Earlier quoted context omitted.
Why not? Fixed-size SIMD architectures use mostly the same operations, so if you target SSE2 initially, the code should run just fine on NEON. A runtime that ships a JIT compiler also has the unique opportunity to further optimize SIMD code by using more lanes or limiting the working set to the host platform's L1 cache size. Even the AOT compilers like GCC or clang emulate platform-specific intrinsics using generic v…
They are similar but not the same, for instance SSE has movemask, but NEON does not, so it gets emulated(slowly) when targeting that platform. The cross lane ops are different enough that you might need to rewrite for other platforms. And then you run into situations where an instruction is very fast on one architecture but horribly slow on another because its basically emulated.
I compared against native: #define ITERATIONS 1000 int main() { const size_t BUFFER_SIZE = 64ul \* 1024 \* 1024; __m128i\* data_buffer = (__m128i *)memalign(64, BUFFER_SIZE); const __m128i all_ones = _mm_set1_epi8(0xFF); for (size_t i = 0; i I had to fixup the WAT because set_local and get_local don't exist anymore. They are called local.get and local.set now. At higher number of iterations the C version converges on…
I compared against native: #define ITERATIONS 1000 int main() { const size_t BUFFER_SIZE = 64ul \* 1024 \* 1024; __m128i\* data_buffer = (__m128i *)memalign(64, BUFFER_SIZE); const __m128i all_ones = _mm_set1_epi8(0xFF); for (size_t i = 0; i I had to fixup the WAT because set_local and get_local don't exist anymore. They are called local.get and local.set now. At higher number of iterations the C version converges on…
I compared against native: #define ITERATIONS 1000 int main() { const size_t BUFFER_SIZE = 64ul \* 1024 \* 1024; __m128i\* data_buffer = (__m128i *)memalign(64, BUFFER_SIZE); const __m128i all_ones = _mm_set1_epi8(0xFF); for (size_t i = 0; i I had to fixup the WAT because set_local and get_local don't exist anymore. They are called local.get and local.set now. At higher number of iterations the C version converges on…
very depressing when for so many use cases even native performance is very very very much not fast enough
If native performance is "very very very" not fast enough then that's supercomputer work and it doesn't really matter if WASM is 3x native or 0.3x native. So that context should be where you're the least depressed.
I compared against native: #define ITERATIONS 1000 int main() { const size_t BUFFER_SIZE = 64ul \* 1024 \* 1024; __m128i\* data_buffer = (__m128i *)memalign(64, BUFFER_SIZE); const __m128i all_ones = _mm_set1_epi8(0xFF); for (size_t i = 0; i I had to fixup the WAT because set_local and get_local don't exist anymore. They are called local.get and local.set now. At higher number of iterations the C version converges on…
very depressing when for so many use cases even native performance is very very very much not fast enough
zig cc -Ofast --target=wasm32-wasi -mcpu=generic+simd128 example.c
zig c++ -Ofast --target=wasm32-wasi -mcpu=generic+simd128 example.cpp
Or with a build file: zig build -Drelease-fast -Dtarget=wasm32-wasi -Dcpu=generic+simd128I compared against native: #define ITERATIONS 1000 int main() { const size_t BUFFER_SIZE = 64ul \* 1024 \* 1024; __m128i\* data_buffer = (__m128i *)memalign(64, BUFFER_SIZE); const __m128i all_ones = _mm_set1_epi8(0xFF); for (size_t i = 0; i I had to fixup the WAT because set_local and get_local don't exist anymore. They are called local.get and local.set now. At higher number of iterations the C version converges on…
Have you tried with the llvm backend? I believe the results might be even better there!
$ time ./wasmer run --llvm fill_buffer.wasm -i fillBufferWithSIMD 1000It looks promising! But fixed-width lanes don't seem too cross-platform? I don't just mean the v256 and v512 types that may become ubiquitous in a few years, but also things like optimizing for different L1 cache sizes, doing some operation macro-fusion on the SIMD unit, or directly supporting leading/trailing elements to reduce code size?
In practice it's not possible to optimize "generally" for all possible target architectures your wasm will run on. You're going to optimize for x86-64 or ARM, and probably going to specifically optimize for modern intel, modern amd, or apple's m1. If you try to optimize for everything you're going to run into really painful tradeoffs and probably have mediocre performance on a bunch of architectures after a lot of ha…
Earlier quoted context omitted.
very depressing when for so many use cases even native performance is very very very much not fast enough
> when If native performance is "very very very" not fast enough then that's supercomputer work and it doesn't really matter if WASM is 3x native or 0.3x native. So that context should be where you're the least depressed.
today's laptop work is late 90's supercomputer's work (and it was even more depressing back then).