Am I hallucinating or didn't several of the examples have background audio artifacts, like it's been trained on speech with noisy backgrounds, I'm guessing audio from movies paired with subtitles? Having random background audio can make it quite hard to use in production.
Bark – Text-prompted generative audio model
31–40 of 99 posts
Re: Bark – Text-prompted generative audio model
#32Am I hallucinating or didn't several of the examples have background audio artifacts, like it's been trained on speech with noisy backgrounds, I'm guessing audio from movies paired with subtitles? Having random background audio can make it quite hard to use in production.
I hear it too. I don't know if it's just background noise though. May be quality issues with the audio synthesis.
Re: Bark – Text-prompted generative audio model
#33Re: Bark – Text-prompted generative audio model
#34Very cool. Side note: bark-gpt.com is already taken for a dog translator: "The world’s first AI powered, real-time communications tool between humans and their furry best friends."[0] I only know this because my law firm partner's name is Bark, and I wanted to automate some legal work and name the software "Bark GPT" after him. [0] https://www.bark-gpt.com/
It's genius idea. As long as you tell owners what they want to believe their pet says, those guys will make a fortune.
Choosing to ignore your dog won't change because some magical AI can now translate it to -Im fine, Im only barking because you're an awesome being, keep your subscription humaaan-
Re: Bark – Text-prompted generative audio model
#35Re: Bark – Text-prompted generative audio model
#36It seems like a lot of the entries in TTS are either close sourced saas apps or something like this with limitations on customizing it. It seems clearly inevitable and likely only months away that a high quality unrestricted open source option for things like voice cloning will emerge so i'm not sure why these projects are even really bothering trying to stop it. I think in order for TTS to have its StableDiffusion m…
CYA aka https://en.wikipedia.org/wiki/Cover_your_ass
also it still requires tons of money to run, so it's likely only businesses will do it
Re: Bark – Text-prompted generative audio model
#37Very cool. Side note: bark-gpt.com is already taken for a dog translator: "The world’s first AI powered, real-time communications tool between humans and their furry best friends."[0] I only know this because my law firm partner's name is Bark, and I wanted to automate some legal work and name the software "Bark GPT" after him. [0] https://www.bark-gpt.com/
Re: Bark – Text-prompted generative audio model
#38Hey, one of the Suno founders/creators of Bark here. Thanks for all the comments, we love seeing how we can improve things in the future. At Suno we work on audio foundation models, creating speech, music, sounds effects etc…. Text to speech was a natural playground for us to share with the community and get some feedback. Given that this model is a full GPT model, the text input is merely a guidance and the model ca…
Amazing work so far! Do you have any sense about how difficult it would be to enable M1/M2 or CoreML support?