> "
his saves on allocations, saves on total memory usage"
Co-Dfns is Aaron Hsu's project, and he has a talk[1] on a compiler implemented in a nano-pass style, once in Racket, once in Chez Scheme[2] and once in Dyalog APL, each using CPU and GPU.
For compiling an AST with ~16 million nodes, he gives benchmark figures that Dyalog APL takes 64MB RAM to hold it, Racket takes 1051MB and Chez Scheme takes 1485MB. His provocative statement for the talk is "pointers are the refined sugar of programming". The APL version is 1/50th of the lines of code[4], runs typically 9x faster on the CPU, 50x-200x faster on the GPU, despite being interpreted by Dyalog vs compiled Scheme. There are benchmarks where the APL version loses in runtime, under 1k lines of code (which he calls "tiny tiny") so lots of startup overhead and not enough big arrays to gain ground on.
He has one talk (available on YouTube) about his approach to using arrays to working with tree structures in a high-performance way in APL[3].
[1] https://www.youtube.com/watch?v=UDqx1afGtQc
[2] he used to be on the Scheme R6RS steering committee, so he's not new/inexperienced with it http://www.r6rs.org/steering-committee/election/electorate.h...
[3] https://www.youtube.com/watch?v=hzPd3umu78g
[4] 17 lines of code which took a PhD researcher (him) 5-10 years to write, mind you.