For inference, if you have a supported card (or probably architecture if you are on Linux and can use HSA_OVERRIDE_GFX_VERSION), then you can probably run anything with (upstream) PyTorch and transformers. Also, compiling llama.cpp is has been pretty trouble-free for me for at least a year. (If you are on Windows, there is usually a win-hip binary of llama.cpp in the project's releases or if things totally refuse to…
has a docker image but no examples to run it https://github.com/ggerganov/llama.cpp/blob/master/docs/dock...
has a docker image but no examples to run it https://github.com/LostRuins/koboldcpp?tab=readme-ov-file#do...
docker image was broken for me on 7800xt running rhel9 https://github.com/Atinoda/text-generation-webui-docker