>” How to prevent this? The obvious way is to just not use dataframes, at least not while doing aggregation. Rather than allocating a huge dataframe and loading our partial results into its columns bit by bit, we can just store our partial results in a plain list.” A lot to be said for not defaulting to data frames, in both r and python. Or, if you must, using something like r data.table or python’s polars if you don…
> A lot to be said for not defaulting to data frames, in both r and python I would even add especially in Python. The main issue I have found is that pandas heavy code is just not as easy to integrate into other Python tools/features/abstractions as code using mostly numpy, dictionaries and various comprehensions to do the vast majority of your work. As a heavy pandas user for several years, I decided about a year ag…
Discovering copy-on-write in R
11–20 of 23 posts
Re: Discovering copy-on-write in R
#12Earlier quoted context omitted.
> A lot to be said for not defaulting to data frames, in both r and python I would even add especially in Python. The main issue I have found is that pandas heavy code is just not as easy to integrate into other Python tools/features/abstractions as code using mostly numpy, dictionaries and various comprehensions to do the vast majority of your work. As a heavy pandas user for several years, I decided about a year ag…
In what cases have you found it worthwhile to use pandas?
(Note that in general, I'm the biggest pandas hater I know)
Re: Discovering copy-on-write in R
#13Good article. Some smaller changes you could make to the final function: In the last line `as.data.frame(do.call(cbind, out_list))` is used to convert the list to a data.frame. Passing it to `cbind` converts the list to a matrix (i.e. combines it into one long vector internally), and then `as.data.frame` converts it back to a list (as noted in the article data frames are lists). Instead, you can use `as.data.frame(ou…
Good advice on `col_grouping` as well, accessing those components of an aggregation rule by index rather than by name is a bad code smell and decreases readability for sure.
Re: Discovering copy-on-write in R
#14>” How to prevent this? The obvious way is to just not use dataframes, at least not while doing aggregation. Rather than allocating a huge dataframe and loading our partial results into its columns bit by bit, we can just store our partial results in a plain list.” A lot to be said for not defaulting to data frames, in both r and python. Or, if you must, using something like r data.table or python’s polars if you don…
> A lot to be said for not defaulting to data frames, in both r and python I would even add especially in Python. The main issue I have found is that pandas heavy code is just not as easy to integrate into other Python tools/features/abstractions as code using mostly numpy, dictionaries and various comprehensions to do the vast majority of your work. As a heavy pandas user for several years, I decided about a year ag…
This is the biggest part. Giving yourself permission to make real abstractions, rather than forcing yourself to go directly from data-on-disk to pandas (or whatever) makes it that much easier to test, repeat, modify, and extend whatever analysis you're working on.
Re: Discovering copy-on-write in R
#15Earlier quoted context omitted.
> A lot to be said for not defaulting to data frames, in both r and python I would even add especially in Python. The main issue I have found is that pandas heavy code is just not as easy to integrate into other Python tools/features/abstractions as code using mostly numpy, dictionaries and various comprehensions to do the vast majority of your work. As a heavy pandas user for several years, I decided about a year ag…
In what cases have you found it worthwhile to use pandas?
Re: Discovering copy-on-write in R
#16https://github.com/smithzvk/modf
With modf, you use the existing place syntax to refer to part of an object. It looks like you're mutating that object, but in fact it will return a clone of the entire containing object, with the modification, while the original remains untouched.
Re: Discovering copy-on-write in R
#17Copy on write is really nice, especially when I often face a very very large read-only matrix (200+GB) and want to do some embarrassingly parallel processes on subsets of it. I haven't found a language which makes it as easy, not python (although not unexpected), not Julia even
If you haven’t tried Swift, copy-on-write is one of the core tenets of its value types (`struct`s, basically), and it’s almost entirely transparent.
Re: Discovering copy-on-write in R
#18Copy on write is really nice, especially when I often face a very very large read-only matrix (200+GB) and want to do some embarrassingly parallel processes on subsets of it. I haven't found a language which makes it as easy, not python (although not unexpected), not Julia even
> not python Pandas has a global option to turn on copy-on-write. https://pandas.pydata.org/docs/dev/user_guide/copy_on_write....
To be default mode in Pandas 3, but seeing as how long it took them to pull the trigger on Pandas 2, that could be a while.
Re: Discovering copy-on-write in R
#19 x
But your z[2] is just an elaborate weighted mean, I would use the following built-in function - weighted.mean(c(2,5),c(1/(1+4),4/(1+4)))
Similarly, your z[3] is just the weighted standard deviation, available in library modi. I was wondering, isn't it better to store the data in some vector x & the weights in a different vector y, and compute the weighted mean & weighted variance in a straightforward fashion like above, or am I missing something. Thanks.Re: Discovering copy-on-write in R
#20Good article. Some smaller changes you could make to the final function: In the last line `as.data.frame(do.call(cbind, out_list))` is used to convert the list to a data.frame. Passing it to `cbind` converts the list to a matrix (i.e. combines it into one long vector internally), and then `as.data.frame` converts it back to a list (as noted in the article data frames are lists). Instead, you can use `as.data.frame(ou…
Thanks for the feedback! The business with `cbind` is a facepalm, I'll definitely fix that. I don't think it will affect performance much since that last step won't be repeated many times, but it makes me cringe now knowing how redundant it is. Good advice on `col_grouping` as well, accessing those components of an aggregation rule by index rather than by name is a bad code smell and decreases readability for sure.