Hey, one of the Suno founders/creators of Bark here. Thanks for all the comments, we love seeing how we can improve things in the future. At Suno we work on audio foundation models, creating speech, music, sounds effects etc…. Text to speech was a natural playground for us to share with the community and get some feedback. Given that this model is a full GPT model, the text input is merely a guidance and the model ca…
This tech will be used by crooks to automate attacks. Generate the language using GPT-4 and the audio using Bark, and then start making phone calls. Because it’s open source, all you need is GPUs. This is not a criticism. I’m impressed and grateful for the openness. Everyone needs to wake up and recognize that these attacks are coming at us essentially right now.
Bark – Text-prompted generative audio model
71–80 of 99 posts
Re: Bark – Text-prompted generative audio model
#72Earlier quoted context omitted.
I already know when my dogs needs to eat, drink, shit or pee or go for a walk or play because he usually will tell me. Choosing to ignore your dog won't change because some magical AI can now translate it to -Im fine, Im only barking because you're an awesome being, keep your subscription humaaan-
But can your dog (translator) say "I love you?" In a doggy-voice? Replika proves people will pay for this and convince themselves it's real, because they want it to be. Bark-GPT's VC pitch: "Replika for real dogs"
Re: Bark – Text-prompted generative audio model
#73Very cool. Side note: bark-gpt.com is already taken for a dog translator: "The world’s first AI powered, real-time communications tool between humans and their furry best friends."[0] I only know this because my law firm partner's name is Bark, and I wanted to automate some legal work and name the software "Bark GPT" after him. [0] https://www.bark-gpt.com/
It's genius idea. As long as you tell owners what they want to believe their pet says, those guys will make a fortune.
Re: Bark – Text-prompted generative audio model
#74Earlier quoted context omitted.
This tech will be used by crooks to automate attacks. Generate the language using GPT-4 and the audio using Bark, and then start making phone calls. Because it’s open source, all you need is GPUs. This is not a criticism. I’m impressed and grateful for the openness. Everyone needs to wake up and recognize that these attacks are coming at us essentially right now.
yawn who cares? If it’s an issue, let law enforcement handle it. Everything has nefarious uses. Humanity marches on.
For reference, look at how the societal damage of social networks has been handled: too little, too late. Same goes for RentTech.
But, I don't know the solution. The common computer has become so powerful that we cannot simply rely on inaccessible materials to prevent the danger of overly-powerful tech spreading too fast, as we do with bioweapons or traditional WMDs.
Fight tech with more tech, I suppose.
Re: Bark – Text-prompted generative audio model
#75I hope you reconsider the misuse mitigation, I'm trying to clone my voice, but so far the other tools weren't that great ...
The misuse mitigation is 3 assert statements in one file. Its pretty easy to comment out.
Re: Bark – Text-prompted generative audio model
#76Re: Bark – Text-prompted generative audio model
#77Earlier quoted context omitted.
This tech will be used by crooks to automate attacks. Generate the language using GPT-4 and the audio using Bark, and then start making phone calls. Because it’s open source, all you need is GPUs. This is not a criticism. I’m impressed and grateful for the openness. Everyone needs to wake up and recognize that these attacks are coming at us essentially right now.
yawn who cares? If it’s an issue, let law enforcement handle it. Everything has nefarious uses. Humanity marches on.
We want to go back to New York in the 90s? Petty theft everywhere?
Re: Bark – Text-prompted generative audio model
#78https://www.youtube.com/watch?v=XSoOrlh5i1k
Someone needs to hook create the plumbing to capture speech to text, feed it to a GPT script that has been told how to reply to such call center calls, then send that back through a TTS generator like this one.
To overcome any latency issues, it could build in a ploy to buy time like the old script did, eg, make the robo-answerer sound like a somewhat addled old man who has to think before each reply, perhaps prefixing responses with "hmm, ahh, ..." to buy time to generate the response.
Re: Bark – Text-prompted generative audio model
#79Earlier quoted context omitted.
Recommend to encode sub-audible tracking symbols in synthesize speech patterns that includes GPS, IP, timestamp, country of origin. We do that already in bootleg movies so we can apply similar methods in synthesizes speech.
It's open source, and even if it wasn't, it would likely be pretty easy to remove those (not that you would get GPS) from an audio file.
Of course, you wouldn't encode real-time navigation data, but a small block of identifying text. Either way, though, someone without a copy of the spreading code isn't going to notice it or decode it. Given enough redundancy in both the time and frequency domains, removing it wouldn't be easy either.
Re: Bark – Text-prompted generative audio model
#80Earlier quoted context omitted.
yawn who cares? If it’s an issue, let law enforcement handle it. Everything has nefarious uses. Humanity marches on.
You'll care when it effects you. We want to go back to New York in the 90s? Petty theft everywhere?