Live data from Hacker News

Bark – Text-prompted generative audio model

github.com

41–50 of 99 posts

Re: Bark – Text-prompted generative audio model

#42
The fact that this is open source and can generate more thann just speech is really nice, but for speech itself, it's much lower quality than what Eleven Labs provides.

All the open source models I've seen so far have this weird kind of neural fuzziness to them. I don't know what Eleven does better, but there's definitely a big difference.

Re: Bark – Text-prompted generative audio model

#43
post #32

Earlier quoted context omitted.

I hear it too. I don't know if it's just background noise though. May be quality issues with the audio synthesis.

yeah sometimes there are definitely artifacts. technically they can be removed pretty easily with another model (like denoiser from FB) but for now we wanted to keep it simple to learn to control these things better through prompt engineering. Like when using a high quality input prompt it generally continues with high quality

At least in the last example, with the man and woman and the expensive oat milk, the background noise seemed to fit a likely public conversation scenario. I wasn't sure if it was accidental or not.

Re: Bark – Text-prompted generative audio model

#45

The fact that this is open source and can generate more thann just speech is really nice, but for speech itself, it's much lower quality than what Eleven Labs provides. All the open source models I've seen so far have this weird kind of neural fuzziness to them. I don't know what Eleven does better, but there's definitely a big difference.

Bark's readme points out that to access the "larger model" you'd have to email them.

I guess, the "open" part of it is mostly for marketing.

Re: Bark – Text-prompted generative audio model

#46
post #34
post #25

Earlier quoted context omitted.

It's genius idea. As long as you tell owners what they want to believe their pet says, those guys will make a fortune.

I already know when my dogs needs to eat, drink, shit or pee or go for a walk or play because he usually will tell me. Choosing to ignore your dog won't change because some magical AI can now translate it to -Im fine, Im only barking because you're an awesome being, keep your subscription humaaan-

But can your dog (translator) say "I love you?" In a doggy-voice? Replika proves people will pay for this and convince themselves it's real, because they want it to be.

Bark-GPT's VC pitch: "Replika for real dogs"

Re: Bark – Text-prompted generative audio model

#49

"However, to mitigate misuse of this technology, we limit the audio history prompts to a limited set of Suno-provided, fully synthetic options to choose from for each language." Isn't this open source and can be easily removed or am I missing something?

Yes, it seems to be enforced by a few assert statements in the code.

Re: Bark – Text-prompted generative audio model

#50
post #23

Hey, one of the Suno founders/creators of Bark here. Thanks for all the comments, we love seeing how we can improve things in the future. At Suno we work on audio foundation models, creating speech, music, sounds effects etc…. Text to speech was a natural playground for us to share with the community and get some feedback. Given that this model is a full GPT model, the text input is merely a guidance and the model ca…

[deleted]
Post reply on HN