However, people here have been spoiled by incredibly good LLMs lately. And the responses that this model gives are nowhere need the high quality of SOTA models today in terms of content. It reminds me more of the 2019 LLMs we saw back in the day.
So I think you've done a "good enough" job on the audio side of things, and further focus should be entirely on the quality of the responses instead.