This is good data, but I'm not sure what the actionable is for me as a Grug Programmer. It means if I'm doing very light processing (sums) I should try to move that to structure-of-arrays to take advantage of cache? But if I'm doing something very expensive, I can leave it as array-of-structures, since the computation will dominate the memory access in Amdahl's Law analysis? This data should tell me something about o…
How much linear memory access is enough?
11–14 of 14 posts
Re: How much linear memory access is enough?
#12Re: How much linear memory access is enough?
#13I looked into this because part of our pipeline is forced to be chunked. Most advice I've seen boils down to "more contiguity = better", but without numbers, or at least not generalizable ones. My concrete tasks will already reach peak performance before 128 kB and I couldn't find pure processing workloads that benefit significantly beyond 1 MB chunk size. Code is linked in the post, it would be nice to see results o…
Re: How much linear memory access is enough?
#14I wonder how much of the cost is coming from the cache misses vs the more frequent indirections/ILP drop? For example, I wonder what this test looks like if you don't randomize the chunks but instead just have the chunks in work order? If you still see the perf hit, that suggests the cost is not from the cache misses but rather the overhead of needing to switch chunks more often.
Note that the base setup has zero cache reuse because each run touches a completely different and cold part of memory. (that makes the result more of an upper bound on the needed chunk size)