Live data from Hacker News

Accents in latent spaces: How AI hears accent strength in English

accent-strength.boldvoice.com

11–20 of 131 posts

Re: Accents in latent spaces: How AI hears accent strength in English

#11
This is super cool.

A suggestion and some surprise: I’m surprised by your assertion that there’s no clustering. I see the representation shows no clustering, and believe you that there is therefore no broad high-dimensional clustering. I also agree that the demo where Victor’s voice moves closer to Eliza’s sounds more native.

But, how can it be that you can show directionality toward “native” without clustering? I would read this as a problem with my embedding, not a feature. Perhaps there are some smaller-dimensional sub-axes that do encode what sort of accent someone has?

Suggestion for the BoldVoice team: if you’d like to go viral, I suggest you dig into American idiolects — two that are hard not to talk about / opine on / retweet are AAVE and Gay male speech (not sure if there’s a more formal name for this, it’s what Wikipedia uses).

I’m in a mixed race family, and we spent a lot of time playing with ChatGPT’s AAVE abilities which have, I think sadly, been completely nerfed over the releases. Chat seems to have no sense of shame when it says speaking like one of my kids is harmful; I imagine the well intentioned OpenAI folks were sort of thinking the opposite when they cut it out. It seems to have a list of “okay” and “bad” idiolects baked in - for instance, it will give you a thick Irish accent, a Boston accent, a NY/Bronx accent, but no Asian/SE Asian accents.

I like the idea of an idiolect-manager, something that could help me move my speech more or less toward a given idiolect. Similarly England is a rich minefield of idiolects, from scouse to highly posh.

I’m guessing you guys are aimed at the call center market based on your demo, but there could be a lot more applications! Voice coaches in Hollywood (the good ones) charge hundreds of dollar per hour, so there’s a valuable if small market out there for much of this. Thanks for the demo and write up. Very cool.

Re: Accents in latent spaces: How AI hears accent strength in English

#13
Victor's problem isn't really the vowels or pacing. The final consonants are soft or not really audible. I am not hearing the /ŋ/ of "long" as the most marked example. It sounds closer to "law". In his "improved" recording he hasn't fixed this.

I sometimes see content on social media encouraging people to sound more native or improve their accent. But IMO it's perfectly ok to have an accent, as long as the speech meets some baseline of intelligibility. (So Victor needs to work on "long" but not "days".) I've even come across people who are trying to mimick a native accent but lose intelligibility, where they'd sound better with their foreign accent. (An example I've seen is a native Spanish speaker trying to imitate the American accent's intervocalic T and D, and I don't understand them. A Spanish /t/ or /d/ would be different from most English language accents, but be way more understandable.)

Re: Accents in latent spaces: How AI hears accent strength in English

#14

This is super cool. A suggestion and some surprise: I’m surprised by your assertion that there’s no clustering. I see the representation shows no clustering, and believe you that there is therefore no broad high-dimensional clustering. I also agree that the demo where Victor’s voice moves closer to Eliza’s sounds more native. But, how can it be that you can show directionality toward “native” without clustering? I wo…

(Minor nitpick, but I think "dialect" is a more appropriate word than "idiolect" here—at least according to Wikipedia, "idiolect" refers to a single person's way of speaking, whereas AAVE et al. are shared and are therefore considered dialects.)

Re: Accents in latent spaces: How AI hears accent strength in English

#18

What a great AI use-case! At first, I felt excited ... But then I read their privacy policy. They want permission to save all of my audio interactions for all eternity. It's so sad that I will never try out their (admittedly super cool) AI tech.

You can reach out and request your data to be deleted at any time.

"if you wish to opt out of future collection of voice samples, you may do so by disabling voice-related features in the BoldVoice app. Please note that this may limit the functionality of certain services."

Yeah, I can opt out. By not using any voice-related feature in their voice training app.

Re: Accents in latent spaces: How AI hears accent strength in English

#20

This is super cool. A suggestion and some surprise: I’m surprised by your assertion that there’s no clustering. I see the representation shows no clustering, and believe you that there is therefore no broad high-dimensional clustering. I also agree that the demo where Victor’s voice moves closer to Eliza’s sounds more native. But, how can it be that you can show directionality toward “native” without clustering? I wo…

> It seems to have a list of “okay” and “bad” idiolects baked in

We're back to "AI safety actually means brand safety": inept pushback against being made into an automated racism factory with their name on it.

Post reply on HN