Live data from Hacker News

Show HN: I trained a 9M speech model to fix my Mandarin tones

simedw.com

61–70 of 166 posts

Re: Show HN: I trained a 9M speech model to fix my Mandarin tones

#61
post #41

Longtime lurker, made an account specifically to give feedback here as an intermediate speaker. :) This is a great initiative and I hope to see more come out of this; I am not criticizing, but just want to provide my user experience here so you have data points. In short, my experience lines up with your native speakers. I found that it loses track of the phonemes when speaking quickly, and tones don't seem to line u…

I don't think it takes care of tone transformation (eg 他是 ni3shi4 -> ni2shi4). Or if it does, my tones are just off. But it's a really cool idea!

[deleted]

Re: Show HN: I trained a 9M speech model to fix my Mandarin tones

#63

How difficult would it be to adapt this to Cantonese? It is a surprisingly difficult language to learn. It has more tones than Mandarin plus comparatively less access to learning resources (in my experience)

Unlike Mandarin and other Chinese languages, Cantonese does not have tone sandhi and has changed tones instead.

Cantonese tones are also different from those of Mandarin, so no, it can't be adopted for Cantonese and it would require a complete rework.

> It is a surprisingly difficult language to learn.

I keep hearing this quite a bit, but I do not find Cantonese to be any more difficult than most languages[0]. Or at least we would need to define a metric based on which we could assess the difficulty. If it is the number of tones, their number (six – no, not nine) may look formidable at first, but they are, in fact, rather simple tones and broadly fall into three categories: flat, rising, and falling. As a random example, Cantonese does not even have a dipping tone.

In comparison, «fancy» tones of Vietnamese are significantly more challenging or even difficult – they can curl and unfurl (so to speak).

[0] That crown appears to belong to Archi, with honourable mentions going out to Inuit, Basque, Georgian, Navajo, Yimas and several other polysynthetic languages.

Re: Show HN: I trained a 9M speech model to fix my Mandarin tones

#64

The article mentions the bitter lesson. I'm confused about the status of Sutton's opinion of the bitter lesson. On the one hand, he invented the concept. On the other hand, he appears to be saying that LLMs are not the correct approach to artificial intelligence, which to a naive outsider looks like a contradiction. What gives?

Maybe he means that LLM will hit a ceiling glass or that the "right" approach will give equivalent with less training/less intensive compute requirements ?

Re: Show HN: I trained a 9M speech model to fix my Mandarin tones

#66
post #51
post #40

Earlier quoted context omitted.

I had the same issue! Perhaps being another dapangzi is the problem here lol

I'm not familiar with this slang: what's a big plate?

It's a slang for somebody fat. 子 does not carry a specific meaning it is more a character with grammatical function to nominative

Re: Show HN: I trained a 9M speech model to fix my Mandarin tones

#67

Earlier quoted context omitted.

In a university Mandarin class, one of the adult students (i.e. probably 40 or so) WAY over exaggerated his tones, to the point that the little old lady teaching us laughed out loud after one of his answers. A few years later, he had the most clean and consistent pronunciation out of anyone I'd been in a class with, and easily switched between the Beijing and other accents depending on which teacher we had on any giv…

From a language learning standpoint that does make sense. Over-exageration while you are learning to help cement the idea, and then when you are speaking more naturally you will fall back into a regular kind of tone.

Over-exaggeration also works well when learning to play stringed instruments like cello.

Re: Show HN: I trained a 9M speech model to fix my Mandarin tones

#70
post #6
post #3

When I was living in Taiwan, one of the ways I forced myself to remember to pronounce the tones distinctly was by waving my hand in front of me, tracing the arc of each character’s tone. It helped a lot even if I did look like an insane expat conducting an invisible orchestra. One more thing: there's quite a bit of variation in how regional accents in the mainland can affect tonal pronunciation. It might be worth rea…

For accents, I’ve mostly tested with a few friends so far. I’m wondering whether region should be a parameter, because training on all dialects might make the system too lax.

Probably be a lot of work but it would be really interesting if you had sufficient data sets to train across accents.

Highly recommend taking a look at Phonemica for this:

https://phonemica.net/

Post reply on HN