Faster Argmin on Floats
algorithmiker.github.io
Faster Argmin on Floats
1–10 of 12 posts
Re: Faster Argmin on Floats
#2Re: Faster Argmin on Floats
#3Re: Faster Argmin on Floats
#4How fast if you write a for loop and keep track of the index and value of the smallest (possibly treating them as ints)?
I wonder could that be made faster by using AVX instructions; they allow to find the minimum value among several u32 values, but not immediately its index.
Re: Faster Argmin on Floats
#5Re: Faster Argmin on Floats
#6I had expected something about algorithms , not Rust-specific implementations .
Re: Faster Argmin on Floats
#7How fast if you write a for loop and keep track of the index and value of the smallest (possibly treating them as ints)?
I hazard to guess that it would be the same, because the compiler would produce a loop out of .iter(), would expose the loop index via .enumerate(), and would keep track of that index in .min_by(). I suppose the lambda would be inlined, maybe even along with comparisons. I wonder could that be made faster by using AVX instructions; they allow to find the minimum value among several u32 values, but not immediately its…
// (initialize ns and idxs by reading from the array
// and adding the apropriate constant to the old value of idxs.)
n_acc = min(n_acc, ns);
const is_new_min = eq(n_acc, ns);
idx_acc = blend(idx_acc, idxs, is_new_min);
Edit: I wrote this with min, eq, blend but you can actually use cmpgt, min, blend to avoid having a dependency chain through all three instructions. I am just used to using min, eq, blend because of working on unsigned values that don't have cmpgtyou can consult the list of toys here: https://www.intel.com/content/www/us/en/docs/intrinsics-guid...
Re: Faster Argmin on Floats
#8How fast if you write a for loop and keep track of the index and value of the smallest (possibly treating them as ints)?
I hazard to guess that it would be the same, because the compiler would produce a loop out of .iter(), would expose the loop index via .enumerate(), and would keep track of that index in .min_by(). I suppose the lambda would be inlined, maybe even along with comparisons. I wonder could that be made faster by using AVX instructions; they allow to find the minimum value among several u32 values, but not immediately its…
e.g. using 4 accumulators instead of 1 accumulator in the naive for loop gives me around a 15%-20% speedup (Not using rust, extremely scalar terrible naive C code via g++ with -funroll-all-loops -march=native -O3)
if we're expressing argmax via the obvious C style naive for loop, or a functional reduce, with a single accumulator, we've forcing a chain dependency that isn't really part of the problem. but if we don't care which argmax-ing index we get (if there are multiple minimal elements in the array) then instead of evaluating the reductions in a single rigid chain bound by a single accumulator, we can break the chain and get our hardware to do more work in parallel, even if we're only single threaded.
anonymoushn is doing something much cleverer again using intrinsics but there's still that idea of "how do we break the dependency chain between different operations so the cpu can kick them off in parallel"
Re: Faster Argmin on Floats
#9How fast if you write a for loop and keep track of the index and value of the smallest (possibly treating them as ints)?
I hazard to guess that it would be the same, because the compiler would produce a loop out of .iter(), would expose the loop index via .enumerate(), and would keep track of that index in .min_by(). I suppose the lambda would be inlined, maybe even along with comparisons. I wonder could that be made faster by using AVX instructions; they allow to find the minimum value among several u32 values, but not immediately its…
Re: Faster Argmin on Floats
#10Earlier quoted context omitted.
I hazard to guess that it would be the same, because the compiler would produce a loop out of .iter(), would expose the loop index via .enumerate(), and would keep track of that index in .min_by(). I suppose the lambda would be inlined, maybe even along with comparisons. I wonder could that be made faster by using AVX instructions; they allow to find the minimum value among several u32 values, but not immediately its…
Yes this is fairly easy to write in AVX, and you can track the index also, honestly the code is cleaner and nicer to read than this mildly obfuscated rust.