Live data from Hacker News

Bark: A transformer based text to audio system

github.com

1–10 of 61 posts

Re: Bark: A transformer based text to audio system

#2
Suno is kind of underselling this with their default voices. With just a little effort you get great stuff, very emotive, fantastic cadences. Here's my female David Attenborough I was playing around yesterday. Not the clearest but charming.

https://user-images.githubusercontent.com/163408/238257231-4...

And Bark is more than TTS. While I haven't had much success with one-shotting full songs with Bark, you build a decent sample library with it. A few recent ones I saved, just randomly concatted here:

https://soundcloud.com/jonathan-fly-620508219/bark-all-night...

Re: Bark: A transformer based text to audio system

#3

Suno is kind of underselling this with their default voices. With just a little effort you get great stuff, very emotive, fantastic cadences. Here's my female David Attenborough I was playing around yesterday. Not the clearest but charming. https://user-images.githubusercontent.com/163408/238257231-4... And Bark is more than TTS. While I haven't had much success with one-shotting full songs with Bark, you build a dec…

Definitely good!

The hardest part still seem to get rid of the synthetic high pitch background noise, I've tried most of the text-to-voice synthesizers and ElevenLabs are still my benchmark.

For example:

https://drive.google.com/file/d/1WL_5gFswhZncERGzSRh2hQ42RVr...

Re: Bark: A transformer based text to audio system

#4
post #3

Suno is kind of underselling this with their default voices. With just a little effort you get great stuff, very emotive, fantastic cadences. Here's my female David Attenborough I was playing around yesterday. Not the clearest but charming. https://user-images.githubusercontent.com/163408/238257231-4... And Bark is more than TTS. While I haven't had much success with one-shotting full songs with Bark, you build a dec…

Definitely good! The hardest part still seem to get rid of the synthetic high pitch background noise, I've tried most of the text-to-voice synthesizers and ElevenLabs are still my benchmark. For example: https://drive.google.com/file/d/1WL_5gFswhZncERGzSRh2hQ42RVr...

I actually think Bark actually beats Eleven right now. Bark tends to add a bit more of a metallic twinge towards the end of longer audio segments (I wonder if this is a bug and not a limitation actually...), but Bark is more expressive.

The trick to exceeding Eleven quality is using multiple speaker prompts, and swapping them in and out over the course of longer texts. Each is a slight variation on the original. If you hand tune the swaps it's super good, but if you just swap in 'beginning of new paragraph' 'continuation speaker' automatically, that alone is a huge boost.

I don't have a clip handy but I will make a YouTube or something this week, because the quality is wild if you do this. There's some passable audio clips on my README here https://github.com/JonathanFly/bark but all those are just using the same speaker for every line. That was all just the first samples I tried, from weeks ago, without really putting any effort into it. You can do a lot better.

The extra expressiveness of Bark does come with it being harder to control and having a bit of a mind of its own at times. (And this can also be very very funny: https://twitter.com/jonathanfly/status/1657658109001596929)

So for a production or real time use case Eleven makes more sense. For example even my best speakers will switch to a new voice mid prompt, once in awhile. (As if the audio clip was from interview segment.)

Re: Bark: A transformer based text to audio system

#5
post #3

Earlier quoted context omitted.

Definitely good! The hardest part still seem to get rid of the synthetic high pitch background noise, I've tried most of the text-to-voice synthesizers and ElevenLabs are still my benchmark. For example: https://drive.google.com/file/d/1WL_5gFswhZncERGzSRh2hQ42RVr...

I actually think Bark actually beats Eleven right now. Bark tends to add a bit more of a metallic twinge towards the end of longer audio segments (I wonder if this is a bug and not a limitation actually...), but Bark is more expressive. The trick to exceeding Eleven quality is using multiple speaker prompts, and swapping them in and out over the course of longer texts. Each is a slight variation on the original. If y…

Thanks, I checked more of the bark examples and they definitely add the "personal" touch to it, elevenlabs sound more stale, so i definitely see what you mean. I also listened a few more times to your sample and it's one of the hardest types of voices to get to sound natural, when i tried a similar on elevenlabs, it sounded a lot worse in terms of the metallic s-s at the end.

I'll set it up this weekend and play around with it :)

Re: Bark: A transformer based text to audio system

#6
post #5

Earlier quoted context omitted.

I actually think Bark actually beats Eleven right now. Bark tends to add a bit more of a metallic twinge towards the end of longer audio segments (I wonder if this is a bug and not a limitation actually...), but Bark is more expressive. The trick to exceeding Eleven quality is using multiple speaker prompts, and swapping them in and out over the course of longer texts. Each is a slight variation on the original. If y…

Thanks, I checked more of the bark examples and they definitely add the "personal" touch to it, elevenlabs sound more stale, so i definitely see what you mean. I also listened a few more times to your sample and it's one of the hardest types of voices to get to sound natural, when i tried a similar on elevenlabs, it sounded a lot worse in terms of the metallic s-s at the end. I'll set it up this weekend and play arou…

>when i tried a similar on elevenlabs, it sounded a lot worse in terms of the metallic s-s at the end.

That's interesting. When I'm judging Bark I'm looking at my own random samples, but for eleven I'm seeing stuff people post on Twitter or YouTube, which I suppose must be cherry-picked. I didn't even realize Eleven did the same thing!

Re: Bark: A transformer based text to audio system

#7

Suno is kind of underselling this with their default voices. With just a little effort you get great stuff, very emotive, fantastic cadences. Here's my female David Attenborough I was playing around yesterday. Not the clearest but charming. https://user-images.githubusercontent.com/163408/238257231-4... And Bark is more than TTS. While I haven't had much success with one-shotting full songs with Bark, you build a dec…

I'm curious, how did you generate the David Attenborough voice? The repo says:

> Bark tries to match the tone, pitch, emotion and prosody of a given preset, but does not currently support custom voice cloning.

Re: Bark: A transformer based text to audio system

#8
post #5

Earlier quoted context omitted.

Thanks, I checked more of the bark examples and they definitely add the "personal" touch to it, elevenlabs sound more stale, so i definitely see what you mean. I also listened a few more times to your sample and it's one of the hardest types of voices to get to sound natural, when i tried a similar on elevenlabs, it sounded a lot worse in terms of the metallic s-s at the end. I'll set it up this weekend and play arou…

>when i tried a similar on elevenlabs, it sounded a lot worse in terms of the metallic s-s at the end. That's interesting. When I'm judging Bark I'm looking at my own random samples, but for eleven I'm seeing stuff people post on Twitter or YouTube, which I suppose must be cherry-picked. I didn't even realize Eleven did the same thing!

I guess it's similar to what you mentioned with bark, the more training and custom adaptation that can be done, the less metallic it will sound. I tested with random voices until i got one that was similar, the existing Bella voice is almost as your sample but has very little metallic s-s.

Re: Bark: A transformer based text to audio system

#9

Suno is kind of underselling this with their default voices. With just a little effort you get great stuff, very emotive, fantastic cadences. Here's my female David Attenborough I was playing around yesterday. Not the clearest but charming. https://user-images.githubusercontent.com/163408/238257231-4... And Bark is more than TTS. While I haven't had much success with one-shotting full songs with Bark, you build a dec…

I'm curious, how did you generate the David Attenborough voice? The repo says: > Bark tries to match the tone, pitch, emotion and prosody of a given preset, but does not currently support custom voice cloning.

>I'm curious, how did you generate the David Attenborough voice? The repo says: >> Bark tries to match the tone, pitch, emotion and prosody of a given preset, but does not currently support custom voice cloning

Check back later in the week, I'll have a bit more on that later after I catch up on actual work and can write a bit.

Post reply on HN