Earlier quoted context omitted.
You're right. In that particular case, ND4J comes to Neanderthal's speed. But only in that particular case; and even then ND4J is still not faster than Neanderthal. My initial quest was to find out whether ND4J can be faster than Neanderthal, and I still couldn't find a case where it is. Although, to my defense, the option in question here is very poorly documented. I've found the ND4J tutorial page where it's mentio…
Fair point we are fixing now: https://github.com/deeplearning4j/deeplearning4j-docs/issues... We will be sending out a doc for this by next week with these updates. Thanks a lot for playing ball here. Beyond that, can you clarify what you mean? Do you mean just the gemm op? For that, that's the only case that mattered for us. We will be documenting the what/how/why of this in our docs. Beyond that, I'm not convinced…
Please also note that Neanderthal also has hundreds of operations. The set of use cases where it scratches itches might be wider and more general than you think.
The reasons I'm showcasing matrix multiplications are:
1. That's what you used in the comparison. 2. It is a good proxy for the overall performance. If matrix multiplication is poor, other operations tend to be even poorer :)
Anyway, as I said, I'll be glad to compare other operations that ND4J excells at, or that anyone think are important.
I would also like to see ND4J's comparisons with Tensorflow or Numpy, or PyTorch, or, JVM based MXNet.