No llama.cpp nor any compilation complexity. Run with two Python commands!
I think you have it backwards. The python (ie, huggingface, etc) implementations of transformers are the complex ones with dependency hell so bad even there's even a layer of package manager / env hell. This version of fastchat (there's 2) required a particular commit of huggingface libs for quite a while. Something that only changed recently. And it'll happen again in the future. Python just hides this complexity...…
also nice to see you again, superkuh. I frequented your IRC channel about a decade ago.