Earlier quoted context omitted.
Cryptographic proof of personhood is going to be a thing, is it not? Outside of BigTech, Signal is as poised as WorldCoin to be just that.
Yes and it's going to be done through digital IDs. Unless something dramatic happens, we're poised to turn to digital IDs linked to your real ID and in turn validating access to apps/communication.
Voicebox: Generative AI model for speech that generalizes across tasks
71–80 of 128 posts
Re: Voicebox: Generative AI model for speech that generalizes across tasks
#72Hopefully it’ll do a LLaMA.
Re: Voicebox: Generative AI model for speech that generalizes across tasks
#73Earlier quoted context omitted.
Because it would be extremely easy to produce it cheaper or record it independently. This opens up for non-signed authors to release audio books.
Yes, but they're not going to pass on costs from not using a voice actor. They're just going to charge what they normally would have, and not worry about having to give so and so a cut.
Re: Voicebox: Generative AI model for speech that generalizes across tasks
#74That little stinger at the end was not as surprising as they thought it was :P It's very cool tech, but it's far from transparent. It has a very obvious "autotune" like sound to it that jumps right out. when they edited that one word it was obvious it had been edited. Again, super cool tech, just not going to replace voice actors or anything.
For me its more like I wake up and check if humans have been replaced yet. Oh good, it's another day that I don't have to share one time pads with my mother to ensure that I'm talking to her and not a simulant performing fraud on a massive scale.
Re: Voicebox: Generative AI model for speech that generalizes across tasks
#75Earlier quoted context omitted.
Just listening to it, it's subjectively not better, but if it's > 10x faster/cheaper, I would use it anyway -- it's good enough to be listenable. Eleven Labs is the first voice synthesis that is good enough that I'd listen to an audiobook generated from it, but pricing is such that it would cost $100 to synthesize a 10 hour audiobook. A little too expensive. If they could get it down to $10 I'd cancel my Audible subs…
Have you tried tortoisetts? I believe eleven labs basically forked that and made improvements on voice quality and speed there
Although it wasn't clear to me how voicebox compares.
Re: Voicebox: Generative AI model for speech that generalizes across tasks
#76Re: Voicebox: Generative AI model for speech that generalizes across tasks
#77Earlier quoted context omitted.
Have you tried tortoisetts? I believe eleven labs basically forked that and made improvements on voice quality and speed there
How does it compare to Voicebox in quality?
1 - Is a real pain to get 'working right' - it's not even remotely batteries included
and, more importantly:
2 - Is incredibly slow. I've been turning Heart Of Darkness into an audiobook as a unit test and it takes ~30m per paragraph, on average. Add to that the occasional hiccup where a block gets transcribed badly (Tortoise occasionally 'drops out' of it's selected voice) and Tortoise only really works if you have a ton of compute and you still don't mind waiting forever.
Re: Voicebox: Generative AI model for speech that generalizes across tasks
#78Earlier quoted context omitted.
Yes and it's going to be done through digital IDs. Unless something dramatic happens, we're poised to turn to digital IDs linked to your real ID and in turn validating access to apps/communication.
And authoritarians everywhere will rejoice (and they will give out the means to duplicate these IDs to a select few in case they need to generate evidence that 'you' have offended the state).
There are alternatives to using the state for this, but they are difficult and fraught with UX issues. Perhaps a decentralised web of trust or some sort of blockchain based registrar of trust that can trace trust routes between mutually distrusting individuals.
Unless such a system is in place, international and strong before states start playing in this space, there isn't much chance of beating a state's approach to the problem.
Just look at https certificates. The current system involves browsers shipping configured to trust a whole bunch of entities I don't really trust, and there has been relatively little interest in trying to build a working decentralised approach to site security.
Re: Voicebox: Generative AI model for speech that generalizes across tasks
#79That little stinger at the end was not as surprising as they thought it was :P It's very cool tech, but it's far from transparent. It has a very obvious "autotune" like sound to it that jumps right out. when they edited that one word it was obvious it had been edited. Again, super cool tech, just not going to replace voice actors or anything.
To me it has a very obvious "Hindi is my native language" accent. I mean after literally the first sentence: "The research team at Meta is excited to share our work...". Ouch. The "our work": just ouch. I was wondering why it wasn't a native english speaker presenting the video when the video is precisely about generating speech.
The first seven seconds are particularly bad.
Don't get me wrong: I've got a lovely french accent when I speak english.
This has either been trained on too many audiobooks spoken by non-natives or they've used their own tech, where the "reference audio" given as input was from a non-native.
In any case something is seriously off.
At 1:59, the "Hi guys, thanks you for tuning in! Today we are going to show you..."... That is obviously an Hindi speaker speaking (it's an example of fixing a real voice by removing background sounds).
I think that the main voice of the video was done by the same person who did the example at 1:59. And I think that they used their example of using a "reference audio".
And that person ain't a native english speaker.
To compare: when the reference audio uses a proper english accent (the example with the "diverse ecosystem" at 0:52), then the output from the text-to-speech sounds native.
I think they just fucked the demo video and it may already be ready for prime time.
Re: Voicebox: Generative AI model for speech that generalizes across tasks
#80I am mostly excited for cheaper audiobooks with consistent voices for different characters.
This is a huge boon for independent authors, until AIs replace us as well :-) .
Things I have learned:
* A good human narrator could do much, much better, but the quality obtained this way is not totally terrible.
* The possibility to produce a section in a matter of minutes is a huge plus. The thing with a book is that it's never totally finished. If you discover a problem after you have submitted your text to a human narrator and paid $ XXXX, there is nothing you can do.
* Currently, there is no platform that I know of distributing and selling books like this. Audible only accepts audiobooks narrated by humans. To my knowledge, platforms that accept ebooks don't handle epub with media overlays. Well, Apple Books say they do but I haven't gotten it to work. There are no alternative platforms for audiobooks that I know of, but I haven't done a ton of research there.
* The possibility to have more control over emotions expressed in the speech could be a bonus, particularly for small, overly dramatic parts of the narration. Coqui TTS new editor is a step in the right direction, but their TTS doesn't sound yet as good as Elevenlabs. Voicebox seems promising, but there is no way to use it at least for now.
* Cost is a big deal 1/3. With my scripts, I pay almost nothing when I fix a typo, since most of the audio is stored in little bits in the database, and only what changes is submitted to the API. But the human time of a narrator costs much more, as it should.
* Cost is a big deal 2/3. As a reader, I have learned that how much a book sells tells me nothing about how much I will like it. But only books that have a potential to sell can afford audiobooks. If I want to listen to a story too quirky to be mainstream, or from an independent author that I follow in Twitter, the chances I'll find it as audiobook are next to none.
* Cost is a big deal 3/3. Voice narration is not the only aspect one needs to pay for. A good story needs an army of editors, proofreaders, and designers. Generally, the more an author or a publisher needs to disburse on those, the more bland and mainstream the book must become to sell and justify the investment.
-----------------------------------
Note that this is a WIP. Book chapter with automatic narration:
An epub with media overlays. It requires an epub reader that supports that standard feature of the epub 3 specification. Currently, and that I know of, there is Thorium and BookFusion for iOS.
https://drive.google.com/file/d/1U8XUB9xhu86JuketGH5WchM0obN...
An MP3 track from the epub above:
https://drive.google.com/file/d/1-u89ee52VZzGZ0oTGC_az5Uqbfs...