Earlier quoted context omitted.
Thank you. The shipped preview has only a bit more than 1500LOC. The VLDB paper was presented at Rio in Aug this year already, but I'll try to come over to LA anyways :)
Karthik, I'm no Spark expert but almost all advice I read is to avoid UDFs if at all possible. Examples below: - https://medium.com/teads-engineering/spark-performance-tunin... - https://www.inovex.de/blog/efficient-udafs-with-pyspark/
There are definitely some differences between the kind of UDFs that Spark supports and the kind that Froid handles. For one, Spark UDFs cannot invoke a Spark SQL query in their definition AFAIK, whereas TSQL functions can. But still, some techniques might be applicable. Definitely worth digging further!