Live data from Hacker News

Voder Speech Synthesizer

griffin.moe

31–40 of 46 posts

Re: Voder Speech Synthesizer

#31

I first heard the Voder as the first sample on the Klatt Record [1]. Unfortunately, there it's credited solely to Homer Dudley; neither Bell Telephone Laboratory nor women like Helen Harper who operated the machine were mentioned. [1]: http://www.festvox.org/history/klatt.html

Odd indeed not to mention the artist playing the sample. Also odd that OP article about the instrument does not mention the inventor.

Re: Voder Speech Synthesizer

#32
post #17

Earlier quoted context omitted.

Probably to some degree, but for your two examples I would argue that isn't necessary: For TTS, the "tone" is something you should encode in the input rather than have TTS figure out. I can imagine ebook > LLM > annotated text with speakers, emotions etc > TTS. So the TTS can remain rather dumb. For the self-driving car, it shouldn't know cultural norms and be "more careful" sometimes. It should always know how much…

> For the self-driving car, it shouldn't know cultural norms and be "more careful" sometimes. It should always know how much it sees and what stoping distance it can get with max breaking and its reaction time and adjust accordingly. I used to live next to two schools. In the morning before school the pavement and road outside my house was always full of school kids on bikes. During this time I'd drive with the assum…

I'd expect self driving cars to have much better sensors and reaction times than we do, and as a consequence not needing to choose between those risks and actually carrying people from one point to another.

But they will probably be way slower than people on streets that are just at the side of sidewalks and full of pedestrians.

Re: Voder Speech Synthesizer

#33
Wolfgang von Kempelen (creator of the fake chess automaton known as the Turk) made a similar thing in the 18th century. [0] It had multiple reeds tuned to the same frequency - conceptually similar to the Voder. It might not be coincidence that Bell Labs developed this, given that Bell himself had also made attempts to improve the design, which is how he ended up inventing the telephone.

[0] https://en.wikipedia.org/wiki/Wolfgang_von_Kempelen%27s_spea...

Re: Voder Speech Synthesizer

#34
post #17

Earlier quoted context omitted.

> For the self-driving car, it shouldn't know cultural norms and be "more careful" sometimes. It should always know how much it sees and what stoping distance it can get with max breaking and its reaction time and adjust accordingly. I used to live next to two schools. In the morning before school the pavement and road outside my house was always full of school kids on bikes. During this time I'd drive with the assum…

I'd expect self driving cars to have much better sensors and reaction times than we do, and as a consequence not needing to choose between those risks and actually carrying people from one point to another. But they will probably be way slower than people on streets that are just at the side of sidewalks and full of pedestrians.

> I'd expect self driving cars to have much better sensors and reaction times than we do

That is never going to happen.

Re: Voder Speech Synthesizer

#35
post #12

This is quite off topic, but it reminded me of something I have been thinking about recently – perhaps at the limit all highly capable narrow AI systems must become generally intelligent. I was thinking about the complexity of expression in TTS voice synthesizers recently and it struck me just how difficult a problem that is. To be as expressive as a human the AI model would need to fully "understand" the context of…

Yes. In Iain M. Banks’s Culture, even the guns are generally intelligent.

Re: Voder Speech Synthesizer

#37
post #20

Earlier quoted context omitted.

I agree that they got pretty good but there’s still something that they get wrong, their intonation is a kind of passable average. If you want to be able to distinguish them from actual human speech pay close attention to intonation/inflection. They’re still very usable, Im not claiming otherwise

I can't hear much difference in the Studio voice: https://cloud.google.com/text-to-speech/docs/wavenet I'm fairly sure I couldn't tell Studio voices and real people apart in a blind test.

It would be a good test. I don't think they are yet indistinguishable, for what it's worth.

Re: Voder Speech Synthesizer

#38

Earlier quoted context omitted.

I made a fork with few more features, it might even work on your phone browser: https://jmiskovic.github.io/voicebox

Thank you for this. I had a lot of fun scaring my cat in bed and it inspired me to become a late middle aged opera savant.

I actually am a late middle-aged opera savant but sadly I have no cat to scare

Re: Voder Speech Synthesizer

#40

Earlier quoted context omitted.

There's also a very nice simulation, where you can play with the very different parts of vocal chords: https://imaginary.github.io/pink-trombone/

I made a fork with few more features, it might even work on your phone browser: https://jmiskovic.github.io/voicebox

Unfortunately this webapp (along with the original Pink Trombone) produces super glitchy audio and consumes 95% of CPU on my Chrome v114.0.5735.198 running on Ubuntu 22.04 (which is running on my Thinkpad X220)
Post reply on HN