Live data from Hacker News

Benchmarking 20 programming languages on N-queens and matrix multiplication

github.com

191–194 of 194 posts

Re: Benchmarking 20 programming languages on N-queens and matrix multiplication

#191

Aren’t JIT languages at a disadvantage since they are benchmarked through the CLI rather than using a benchmarking library to allow JIT to warmup?

Some language implementations have run time startup costs (bytecode verifier), not-required at run time by other language implementations.

https://www.oracle.com/java/technologies/security-in-java.ht...

Why should that run time startup cost be ignored?

Re: Benchmarking 20 programming languages on N-queens and matrix multiplication

#192

Doesn’t include Prolog which has decent constraint solver answers for n-queens and sudoku, which are pretty fast but I don’t know how they would compare in benchmarks: https://www.metalevel.at/queens/ https://www.metalevel.at/sudoku/

Make a pull request.

Re: Benchmarking 20 programming languages on N-queens and matrix multiplication

#193

Earlier quoted context omitted.

No, multi-threaded OpenBLAS improves performance to 0.15s.

I dunno man. My claim was that for specific cases with unique properties, it's not hard to beat BLAS, without getting too exotic with your code. BLAS doesn't have routines for multiplies with non-contiguous data, various patterns of sparsity, mixed precision inputs/outputs, etc. The example I gave is for a specific case close-ish to the case I cared about. You're changing it to a very different case, presumably one t…

When I run your benchmark with matrices larger than 1024x1024 it errors out in the verification step. Since your implementation isn't even correct I think my original point about OpenBLAS being extremely difficult to replicate still stands.

Re: Benchmarking 20 programming languages on N-queens and matrix multiplication

#194

Earlier quoted context omitted.

The compilation command errors out for me: /home/bjourne/p/laser/benchmarks/gemm/gemm_bench_float32.nim(77, 8) Warning: use `std/os` instead; ospaths is deprecated [Deprecated] /home/bjourne/p/laser/benchmarks/gemm/gemm_bench_float32.nim(101, 8) template/generic instantiation of `bench` from here /home/bjourne/p/laser/benchmarks/gemm/gemm_bench_float32.nim(106, 21) template/generic instantiation of `gemm_nn_fallback`…

Ah, It was from an older implementation that wasn't compatible with Nim v2. I've commented it out. If you pull again it should work. > Anyway the reason for your competitive performance is likely that you are benchmarking with very small matrices. OpenBLAS spends some time preprocessing the tiles which doesn't really pay off until they become really huge. I don't get why you think it's impossible to reach BLAS speed.…

Ok, I will benchmark this more when I have time. My gut feeling is that it is impossible to reach OpenBLAS-like performance without carefully tuned kernels and without explicit SIMD code. Clearly, it's not impossible to be as fast as OpenBLAS otherwise OpenBLAS itself woudln't be that fast, but it is very difficult and takes a lot of work. There is a reason much of OpenBLAS is implemented in assembly and not C.
Post reply on HN