I can remember the columnar approach used in Fortran back in the day. Fortran did not have records, only scalar arrays, so a natural way to represent a bunch of "objects" was to have a bunch of arrays for each property value. IDK if such prior art is any helpful today, though :)
That being said, statistical computing is at the mercy of the high cost of matrix multiplication.
I think that you can alter the order of arrays in numpy, and examine the performance difference.