Live data from Hacker News

Bark – Text-prompted generative audio model

github.com

71–80 of 99 posts

Re: Bark – Text-prompted generative audio model

#71
post #59
post #23

Hey, one of the Suno founders/creators of Bark here. Thanks for all the comments, we love seeing how we can improve things in the future. At Suno we work on audio foundation models, creating speech, music, sounds effects etc…. Text to speech was a natural playground for us to share with the community and get some feedback. Given that this model is a full GPT model, the text input is merely a guidance and the model ca…

This tech will be used by crooks to automate attacks. Generate the language using GPT-4 and the audio using Bark, and then start making phone calls. Because it’s open source, all you need is GPUs. This is not a criticism. I’m impressed and grateful for the openness. Everyone needs to wake up and recognize that these attacks are coming at us essentially right now.

yawn who cares? If it’s an issue, let law enforcement handle it. Everything has nefarious uses. Humanity marches on.

Re: Bark – Text-prompted generative audio model

#72
post #34

Earlier quoted context omitted.

I already know when my dogs needs to eat, drink, shit or pee or go for a walk or play because he usually will tell me. Choosing to ignore your dog won't change because some magical AI can now translate it to -Im fine, Im only barking because you're an awesome being, keep your subscription humaaan-

But can your dog (translator) say "I love you?" In a doggy-voice? Replika proves people will pay for this and convince themselves it's real, because they want it to be. Bark-GPT's VC pitch: "Replika for real dogs"

[deleted]

Re: Bark – Text-prompted generative audio model

#73
post #25
post #7

Very cool. Side note: bark-gpt.com is already taken for a dog translator: "The world’s first AI powered, real-time communications tool between humans and their furry best friends."[0] I only know this because my law firm partner's name is Bark, and I wanted to automate some legal work and name the software "Bark GPT" after him. [0] https://www.bark-gpt.com/

It's genius idea. As long as you tell owners what they want to believe their pet says, those guys will make a fortune.

your app is in a bad place when fraud is a part of your pitch don’t you think? …

Re: Bark – Text-prompted generative audio model

#74
post #71
post #59

Earlier quoted context omitted.

This tech will be used by crooks to automate attacks. Generate the language using GPT-4 and the audio using Bark, and then start making phone calls. Because it’s open source, all you need is GPUs. This is not a criticism. I’m impressed and grateful for the openness. Everyone needs to wake up and recognize that these attacks are coming at us essentially right now.

yawn who cares? If it’s an issue, let law enforcement handle it. Everything has nefarious uses. Humanity marches on.

As technology gets more powerful more quickly, and the rule of law becomes more and more unable to prevent societal damage, this response becomes woefully inadequate.

For reference, look at how the societal damage of social networks has been handled: too little, too late. Same goes for RentTech.

But, I don't know the solution. The common computer has become so powerful that we cannot simply rely on inaccessible materials to prevent the danger of overly-powerful tech spreading too fast, as we do with bioweapons or traditional WMDs.

Fight tech with more tech, I suppose.

Re: Bark – Text-prompted generative audio model

#75
post #33

I hope you reconsider the misuse mitigation, I'm trying to clone my voice, but so far the other tools weren't that great ...

The misuse mitigation is 3 assert statements in one file. Its pretty easy to comment out.

Slightly trickier is how to generate a "history" file that will allow the model to use your voice.

Re: Bark – Text-prompted generative audio model

#77
post #71
post #59

Earlier quoted context omitted.

This tech will be used by crooks to automate attacks. Generate the language using GPT-4 and the audio using Bark, and then start making phone calls. Because it’s open source, all you need is GPUs. This is not a criticism. I’m impressed and grateful for the openness. Everyone needs to wake up and recognize that these attacks are coming at us essentially right now.

yawn who cares? If it’s an issue, let law enforcement handle it. Everything has nefarious uses. Humanity marches on.

You'll care when it effects you.

We want to go back to New York in the 90s? Petty theft everywhere?

Re: Bark – Text-prompted generative audio model

#78
A few years a back someone had the genius idea of making a robo-answering phone program that would give vague but encouraging replies when it received some unsolicited sales pitch. Although it was a fixed sequence of responses, it fooled some callers for a surprisingly long time.

https://www.youtube.com/watch?v=XSoOrlh5i1k

Someone needs to hook create the plumbing to capture speech to text, feed it to a GPT script that has been told how to reply to such call center calls, then send that back through a TTS generator like this one.

To overcome any latency issues, it could build in a ploy to buy time like the old script did, eg, make the robo-answerer sound like a somewhat addled old man who has to think before each reply, perhaps prefixing responses with "hmm, ahh, ..." to buy time to generate the response.

Re: Bark – Text-prompted generative audio model

#79
post #68

Earlier quoted context omitted.

Recommend to encode sub-audible tracking symbols in synthesize speech patterns that includes GPS, IP, timestamp, country of origin. We do that already in bootleg movies so we can apply similar methods in synthesizes speech.

It's open source, and even if it wasn't, it would likely be pretty easy to remove those (not that you would get GPS) from an audio file.

Hmm, I don't know about that. The data rate used by civilian GPS (L1 C/A) is only 50 bps. The symbols are normally spread over a couple MHz of bandwidth to make it possible to recover at levels below the thermal noise floor. I see no reason why the same thing couldn't be done at baseband, adding an imperceptible bit of extra noise to an audio signal.

Of course, you wouldn't encode real-time navigation data, but a small block of identifying text. Either way, though, someone without a copy of the spreading code isn't going to notice it or decode it. Given enough redundancy in both the time and frequency domains, removing it wouldn't be easy either.

Re: Bark – Text-prompted generative audio model

#80
post #71

Earlier quoted context omitted.

yawn who cares? If it’s an issue, let law enforcement handle it. Everything has nefarious uses. Humanity marches on.

You'll care when it effects you. We want to go back to New York in the 90s? Petty theft everywhere?

The technology already exists, shutting down a company and pretending that it doesn't doesn't solve the problem.
Post reply on HN