Just today, I finished a blog post (also my latest submission, felt like could be useful to some) about how to get something like this working in a bundle of something to run models, as well as a web UI for more easy interaction - in my case that was koboldcpp, which can run GGML, both on the CPU (with OpenBLAS) and on the GPU (with CLBlast). Thanks to Hugging Face, getting Metharme, WizardLM or other models is also…
e.g.
https://replicate.com/stability-ai/stablelm-tuned-alpha-7b
https://github.com/runpod/serverless-workers/tree/main/worke...