Live data from Hacker News

The most frequent 777 characters give 90% coverage of Kanji in the wild

japanesecomplete.com

211–214 of 214 posts

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#211
post #164

Earlier quoted context omitted.

No one knows "all Kanji" much like no one knows "all English words" even in native English-speaking countries.

Do kanji handwriting-recognition models even “know” (i.e. are trained on) all in-use kanji, or do they just not bother with some of the rarest ones because their presence in the dataset would decrease the likelihood of correctly matching more-common characters, which it’s overwhelmingly-likely are what you’re writing?

Presumably it's the same as words in English, where you'd never make a word suggestion model that just blindly considers all possible words. You swipe jow or gow and it will change it to "how" despite the other two existing in arguably correct English in some very obscure case.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#212

Earlier quoted context omitted.

I don't know for Japanese as meaning sometimes shifts from Chinese, but in Chinese the standard definition of 可 is "can, may, be able to". You obviously learn it by itself but as Chinese words are mostly a combination of 2 characters, you immediately also have to learn e.g. 可以 (can, may, be able to), 可能 (maybe), 可爱 (cute) etc. So someone who's learning characters in order to get 90% coverage (or whatever) would not s…

> the meaning of 可爱 would be fairly straightforward to guess To be honest this whole thread about 可愛い is more or less bonkers, because it's an ateji. The word's meaning doesn't derive from the characters, the characters got arbitrarily attached to an existing word because they were similar in sound and meaning. As such, the whole thing is about as meaningful as talking about how easy it is to guess that 珈琲 means "cof…

If the characters were chosen because they fit the meaning, what does that change?

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#213

Earlier quoted context omitted.

> the meaning of 可爱 would be fairly straightforward to guess To be honest this whole thread about 可愛い is more or less bonkers, because it's an ateji. The word's meaning doesn't derive from the characters, the characters got arbitrarily attached to an existing word because they were similar in sound and meaning. As such, the whole thing is about as meaningful as talking about how easy it is to guess that 珈琲 means "cof…

If the characters were chosen because they fit the meaning, what does that change?

The thread was about how words' characters and meanings relate, and with ateji that relation is highly atypical, is the basic point.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#214
post #45
post #40

Earlier quoted context omitted.

This is closer to 90:25 rather than the typical Pareto of 80:20

80:20 is a pop-culture take on the distribution. It does a good job of visualising it to someone who doesn't know maths. It's really any distribution with the CDF of the form x^(-a)

The wikipedia article says it's commonly formulated as 80:20, so that's where I'm getting my info. You're saying it covers every nice Pareto ratio, which is very different. Because 90:25 is much better than 80:20 but they are part of the same phenomenon. Well, you can call everything a Pareto phenomenon then and what's the point if everything fits in this universalish category? How can I explain the value of 90:25 without invoking Pareto and having it constantly diluted to 80:20?
Post reply on HN