Live data from Hacker News

Road to NumPy 2.0

hackmd.io

21–24 of 24 posts

Re: Road to NumPy 2.0

#21
post #7

Still no NA?

I’m working on adding missing data support for strings as part of adding a UTF-8 variable-width string type to NumPy. Not a general solution but should help with a lot of use-cases. https://numpy.org/neps/nep-0055-string_dtype.html

The current memory use of string arrays is another major issue, glad to see this being worked on!

Re: Road to NumPy 2.0

#22

Earlier quoted context omitted.

Why not use a structured array with an 'isna' field to use as a mask when performing operations?

How is that convenient? Missing data support belongs deep in NumPy itself (or any other similar package) so that operations can do the right thing and missing values propagate correctly. For example, let's say you want by definition missing values to sort last. If you roll out your custom missing value marker, you'll also need to roll out your own custom sort function. And the same for a whole lot more stuff.

What about a MaskedArray? ndarrays are homogeneous by definition.

Re: Road to NumPy 2.0

#23
post #16
post #4

Earlier quoted context omitted.

Just FYI, numpy by no means pioneered this concept. There was Fortran before, as a language built around n-dimensional arrays. And even Fortran was not the first one, as some stack based programming languages (such as APL, IIRC) have similar concepts. Also languages such as R and evventually Matlab (as a kind-of nicer frontend to Fortran libs) pioneered this concept. However, Numpy was the first library bringing this…

Numeric Python, the original numpy, was based mainly on the MATLAB array object. Of course, it was also influenced by FORTRAN and other languages, but ultimately, it was taken from the mental model of MATLAB.

and MATLAB was entirely influenced by Fortran's array-based syntax.

Re: Road to NumPy 2.0

#24
post #4

I think that ndarray is the most successful abstraction I've come across. Numerical computing is a domain ripe for terrible code, but multi-indexing, broadcasting, mask arrays, .reshape(), .where(), linspace(), and all that are so well made and useful they are now the standard grammar of data science. Yes you've seen horrible numpy code before, but how much worse would it be if it had been written by the same person…

Just FYI, numpy by no means pioneered this concept. There was Fortran before, as a language built around n-dimensional arrays. And even Fortran was not the first one, as some stack based programming languages (such as APL, IIRC) have similar concepts. Also languages such as R and evventually Matlab (as a kind-of nicer frontend to Fortran libs) pioneered this concept. However, Numpy was the first library bringing this…

Fortran was created in the 1950s. APL — 60s.
Post reply on HN