Earlier quoted context omitted.
Yeah, as long as it’s intelligible an accent is perfectly fine It’s also perfectly fine to want to sound like a native speaker - whether it be because they are self conscious, think it will benefit them in some way, or simply want to feel like they are speaking “correctly” Sorry to pick on you, it’s just amazing to me how sensitive we are to “inclusivity” to the point where we almost discourage people wanting to fit…
That kind of implies that there's a "correct" accent for English, even though many countries and regions natively speak it. Someone from Glasgow is just as much of a native speaker as someone from Los Angeles even though the accents are wildly different.
Accents in latent spaces: How AI hears accent strength in English
91–100 of 131 posts
Re: Accents in latent spaces: How AI hears accent strength in English
#92Like others recently, I've been extremely impressed by LLM's ability to play GeoGuessr, or, more generally, to geo-locate random snapshots that you give them, with what seem (to me) to be almost no context clues. (I gave ChatGPT loads of holiday snapshots, screenshotted to remove metadata, and it did amazingly.) I assume that, with enough training, we could get similarly accurate guesses of a person's linguistic hist…
We actually did something like this for non-native English speakers a few months back. Check out https://accentoracle.com (most mind-blowing if you're a non native English speaker)
Re: Accents in latent spaces: How AI hears accent strength in English
#93Re: Accents in latent spaces: How AI hears accent strength in English
#94On the one hand, the tech is impressive, and the demo is nicely done.
On the other, I think the demo completely misses the point. There's a disconnect between what learners need to learn and what this model optimises for, and it's probably largely explainable by how difficult (maybe even impossible) getting training datasets is. That, and marketing.
I believe most learners optimise for two things: being understood [1] and not being grating to the ear [2]. Both goals hinge on acquiring the right set of phonemes and phonetic "tools", because the sets of meaningfully distinct sounds (phonemes) and "tools" rarely match between languages.
For example, most (all?) Slavic languages have way fewer meaningfully distinct vowels than English. Meaningfully distinct is the crucial part. Russian word "молоко" as it's most often pronounced has three different vowels, at least two of which would be distinct to an English speaker, but Russian speakers hear that as one-ish vowel. And I mean "hear it": it's not a conscious phenomenon! Phoneme recognition is completely subconscious, so unless specifically trained, people often don't hear the difference between sounds that are obviously different to people who speak that language natively [3].
Same goes for phonetic "tools". English speakers shorten vowels when followed by non-voiced consonants, which makes "heart" and "hard" distinguishable even when t/d are transformed into the same sound (glottal stop or a tap). This "tool" is not available in many languages, so people use it incorrectly and it sounds confusing.
So, how would ML models learn this mapping between sounds and phonemes, especially when it's non-local (like with the preceding vowel length)? It's relatively straightforward to find large sets of speech samples labelled with their speakers' backgrounds, but that's just sounds, not phonemes! There is very little signal showing which sound structures matter for humans listening to the sound and which don't. [4]
There's also a set of moral issues connected to the "target accent" approach. Teaching learners to acquire an accent that superficially sounds like whatever they chose as a "target" devalues all other accents, that are just as valid and are just as English, because they have the same phonetic system (phonemes + "tools"). It can also make people sound a bit cringe, which I saw first hand.
Ideally learners should learn phonetic systems, not superficial accents. That's what makes speech intelligible and natural, even if it's has an exotic flavour [5][6]. Systems like the one the company is building do the opposite. I guess they are easier to build and easier to sell.
[1]: On that path lies a nice surprise: being understood and understanding are two sides of the same medal, so learning how to be understood a language learner inevitably starts to understand better. Being able to hear the full set of phonemes is the key to both.
[2]: There's a vast, VAST difference between people not paying attention to how someone speaks and them not being able to tell that something's off when prompted.
[3]: Nice demonstration for non-Hindi speakers: https://www.youtube.com/watch?v=-I7iUUp-cX8 When isolated and spoken slowly, the consonants might sound different, but in normal speech they sound practically indistinguishable to English speakers with no prior exposure. Native speakers would hear the difference as clear as you would in cap/cup!
[4]: Take their viral accent recognition demo. Anecdotally, among three non-native speakers with different backgrounds I talked to, the demo was guessing the mother tongue much better than native speakers, and it errors were different. This is a sign of the model learning to recognise the wrong things.
[5]: Ever noticed how films almost always cast native English speakers imitating non-English accents rather than people for whom that's their first language? That's why, English phonetic system with sprinkles of phonetic flavour is much more understandable.
[6]: By the way, Arnold Schwarzenegger understands this very well.
Re: Accents in latent spaces: How AI hears accent strength in English
#95This is so cool. Real-time accent feedback is something language learners have never had throughout all of human history, until now. Along similar lines, it would be useful to map a speaker's vowels in vowel-space (and likewise for consonants?) to compare native to non-native speakers. I can't wait until something like this is available for Japanese.
Do you have a source for this? It doesn't seem plausible to me, but I'm not an expert.
Re: Accents in latent spaces: How AI hears accent strength in English
#96What a great AI use-case! At first, I felt excited ... But then I read their privacy policy. They want permission to save all of my audio interactions for all eternity. It's so sad that I will never try out their (admittedly super cool) AI tech.
Re: Accents in latent spaces: How AI hears accent strength in English
#97Victor's problem isn't really the vowels or pacing. The final consonants are soft or not really audible. I am not hearing the /ŋ/ of "long" as the most marked example. It sounds closer to "law". In his "improved" recording he hasn't fixed this. I sometimes see content on social media encouraging people to sound more native or improve their accent. But IMO it's perfectly ok to have an accent, as long as the speech mee…
A very important part of people trusting you is them being able to understand what you say without making extra efforts compared to a native speaker. An easy way to improve intonation and fluency is to imitate a native speaker. Copying things like the intervocalic T and D is a consequence of that. It would be easier for a native Spanish speaker to say the Spanish /t/ and /d/ but intonation and fluency would be impair…
Re: Accents in latent spaces: How AI hears accent strength in English
#98Earlier quoted context omitted.
That kind of implies that there's a "correct" accent for English, even though many countries and regions natively speak it. Someone from Glasgow is just as much of a native speaker as someone from Los Angeles even though the accents are wildly different.
Hence in quotes, man
And I've heard other such stories of American schools flagging kids for speech therapy when what they have is an accent. I feel like Americans are actually some of the worst about that.
Re: Accents in latent spaces: How AI hears accent strength in English
#99Earlier quoted context omitted.
Hence in quotes, man
I've literally heard a story of a kid arriving to the US from Scotland and being sent to speech therapy to remove a Scottish accent. And I've heard other such stories of American schools flagging kids for speech therapy when what they have is an accent. I feel like Americans are actually some of the worst about that.
Besides it’s not like there isn’t correct either - if you’re out in the Midwest what’s correct is just what everyone is speaking.
Its obvious that a kid from Ohio who speaks perfect isn’t going to go to Scotland and speak it “correctly”
Like it’s such low hanging fruit to always be that guy to point out the lowest level, most obvious exception
Re: Accents in latent spaces: How AI hears accent strength in English
#100This is so cool. Real-time accent feedback is something language learners have never had throughout all of human history, until now. Along similar lines, it would be useful to map a speaker's vowels in vowel-space (and likewise for consonants?) to compare native to non-native speakers. I can't wait until something like this is available for Japanese.
A good accent coach would be able to do much better by identifying exactly how you're pronouncing things differently, telling you what you should be doing in your mouth to change that, and giving you targeted exercises to practice.
Presumably a model that predicts the position of various articulators at every timestamp in a recording could be useful for something similar.