This makes me wonder: what’s the status of GPU programming on the JVM? Any abstraction for GPGPU or shaders programming?
But it's a research project.
11–19 of 19 posts
This makes me wonder: what’s the status of GPU programming on the JVM? Any abstraction for GPGPU or shaders programming?
But it's a research project.
A Java port of llama2.c that performs very close to C on large models. Llama 2 7B runs at a whooping 1.6 tokens/s.
The Java code is impressively written, using newer features like MemorySegment. Looked at the author and realized it's Alfonso from the Graal team -- makes sense. I wonder whether the "matmul" code could be further optimized with the Vector API and SIMD.
A Java port of llama2.c that performs very close to C on large models. Llama 2 7B runs at a whooping 1.6 tokens/s.
Hey man, awesome stuff. Surely any JIT compiler will struggle to vectorize something using IntStream.range, though? Looking at matmul, I'd not expect that to be auto-vectorized. The Panama API can be used to do a matmul vectorization, too bad it seems to never launch.
Is there any indication that it won't go from there to a final release soon?
Earlier quoted context omitted.
I know it might be asking a lot, but it would be great if someone could put up a HF space so I could try all the various flavours/sizes.
/r/LocalLLaMA/
This makes me wonder: what’s the status of GPU programming on the JVM? Any abstraction for GPGPU or shaders programming?
Earlier quoted context omitted.
Hey man, awesome stuff. Surely any JIT compiler will struggle to vectorize something using IntStream.range, though? Looking at matmul, I'd not expect that to be auto-vectorized. The Panama API can be used to do a matmul vectorization, too bad it seems to never launch.
Panama is now in its third preview in the soon-to-be-released JDK 21: https://openjdk.org/jeps/442 Is there any indication that it won't go from there to a final release soon?
This makes me wonder: what’s the status of GPU programming on the JVM? Any abstraction for GPGPU or shaders programming?
Just in case if anyone interested in Python version, I spend some time on weekend and ported it to pure python -- https://github.com/tairov/llama2.py
I never knew that it would take about 500 lines of core part code to implement inference for such a cutting edge AI technology.