I've always wondered if there's a better way of making voice assistants. With this stack, the AI will not be able to answer "what is this sound?", or give you UK-based information because it picked up on your British accent. It's bottlenecked by text. A model that can understand audio as input, and output audio directly, could be so much more powerful
Sure there's a better way. https://google-research.github.io/seanet/audiopalm/examples/ There's no reason autoregressive LMs can't be used to model audio data.
Re: Show HN: Gdańsk AI – full stack AI voice chatbot
#31Unfortunately this hasn't been released yet and since that paper it is very quiet around that topic so chances are high it (or something similar) will just not be released.