Live data from Hacker News

The most frequent 777 characters give 90% coverage of Kanji in the wild

japanesecomplete.com

181–190 of 214 posts

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#181
post #173

A simple information-theoretic argument suggests that most of the information is in the remaining article 10%.

Not enough information. For example, in English some of the least common characters are also generally not important for conveying information. Your argument only holds if Kanji was designed to be optimal in compressing information, which is of course not how Kanji came to be. That doesn't mean you are wrong, but your argument is invalid.

Aren't kanji closer to english words than english letters?

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#182

A few years ago I downloaded several hundreds of megabytes of Japanese subtitles, split into 3 categories: live action/drama, anime and foreign film/tv I’ve listed them in a google sheets together with a few other corpora https://docs.google.com/spreadsheets/d/1yb5dq4ahdwc_g0aQTL3Y... Choose the jimaku tab for subtitles to see how big the variation between corpus can be. According to other comments here, it appears t…

The source links appear to no longer work. Do you know where we can download Japanese subtitles? I would love to attempt to segment a bunch of Japanese subtitles into words and then do frequency analysis. My interest is in increasing my listening ability, so I want to put the most frequently spoken words into SRS/Anki, and perhaps even break it down by anime. Alternatively, has anyone already done this?

That was my initial goal, but I had a lot of trouble with vanilla MeCab not understanding a lot of the text. But this was before neologd, so i think it would work better now.

I don’t have the source code on me, but I scraped it from a website that publishes subtitles. The scraping was easy, the cleaning not, and I believe this spreadsheet is generated from my first attempt at cleaning.

A lot of sources in Japanese nlp and linguistics have a bad habit of changing url often, so it bitrots easily. Sorry.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#183

Earlier quoted context omitted.

It's no harder than having conversations in spoken Japanese where you don't see any text, kanji or otherwise. It would look awkward initially, but that's just because people are just used to the status quo which is kanji-kana mix.

Spoken Japanese has intonation and rhythm which make identifying word boundaries and homonyms easier. Reading hiragana (especially without space as word boundaries) is totally different. Reading long hiragana by speaking aloud help, but there is still a problem of ha/wa and he/e.

Using space is a given if you write exclusively in kana. See Kanamoji-kai:

http://kanamozi.org/

Also, the very reason there are so many homonyms in written text in the first place is because of the kanji (over)usage - that is, because people think there are visual cues they are less careful about choosing words that are also understood easily by people listening to the words.

When people speak, at least if they are a competent speaker, people tend to avoid the overuse of homonyms (mostly kango).

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#184

If Japanese people just used kana (like Korean people use hangul today), kids in Japan don't have to spend countless hours learning this complex (and often irrational) writing system. http://kanamozi.org/hikari959-04.html > If you really want to be native-level Japanese, kanji are essential There are visually impaired people who have difficulty learning kanji but speak Japanese fluently. Language is not just for peop…

The good news is that this is a problem that's solving itself. Kanji, especially complex kanji, are dying out in favour of katakana. An anecdotal example, as little as 20 years ago, all the fish markets had signs for each fish, in kanji. Now all of those signs are in katakana. Many young people can't read the kanji for specific fish anymore. They can probably manage tuna, but hake? Pollock? Most written communication…

I think it's good that more private businesses are using kana-based writing (including more widespread furigana usage), but I don't consider the problem being fixed until the mandatory teaching of kanji in compulsory education stops, and maybe I just don't know but there's no sign of that happening.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#185
post #95

Earlier quoted context omitted.

> Even if you "know" all Kanji in a word you'll likely not understand the word's meaning unless it's something simple and concrete. I don't think anyone passingly familiar with Japanese thinks otherwise. That is, I think you're arguing against a position here that nobody actually holds. The argument for memorizing kanji is that it makes it easier to learn compound words, not that you'll just know them without learnin…

I agree that people with some Japanese knowledge likely already know this. But out there you will find a lot of marketing material/posts targeting absolute beginners promoting some kind of "Top X% Kanji lists", as if they are a huge shortcut and secret to quickly learning Japanese. So I'm just saying they aren't. If you talk to people who don't know much about Japanese they often believe that memorizing the character…

I feel like a lot of this perception comes from a tendency for people to equate Kanji characters with words. I always explain to others that a character is kind of like a root, e.g. Sub-optimus-al = suboptimal. It would be crazy to say one can learn English by memorizing just a few hundred Latin roots and suffixes, and in fact such a claim is so irrelevant nobody even keeps track of the statistics.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#186
post #6

The thing is, 90% coverage is not that great. What happens is that you understand common words that make up for a lot of structure, but when an uncommon word appears it's probably important to the sentence. For example, "son, if you go to the plumbf tomorrow morning don't forget to pick up some zlonks." 98% is closer to what you need in order to read a text and have an idea of what's going on. See this article: https…

[deleted]

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#187
post #6

The thing is, 90% coverage is not that great. What happens is that you understand common words that make up for a lot of structure, but when an uncommon word appears it's probably important to the sentence. For example, "son, if you go to the plumbf tomorrow morning don't forget to pick up some zlonks." 98% is closer to what you need in order to read a text and have an idea of what's going on. See this article: https…

I laughed so hard at your great example and read it to all my friends sitting in the room.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#188
post #12
post #6

The thing is, 90% coverage is not that great. What happens is that you understand common words that make up for a lot of structure, but when an uncommon word appears it's probably important to the sentence. For example, "son, if you go to the plumbf tomorrow morning don't forget to pick up some zlonks." 98% is closer to what you need in order to read a text and have an idea of what's going on. See this article: https…

As someone who barely scraped by the Kanji/readings of my N2 but have to do a chunk of my work in Japanese, I gotta disagree. Sure, reading isolated sentences with only 90% coverage is really hard sometimes. But usually you're reading whole blocks of text, so you have context. Also, maybe you don't know some kanji but it has similar radicals to other, so you can make educated guesses (only sometimes of course). I mea…

I agree with rtpg on this one - and disagree with the Diego.

I also read/type Japanese (my writing out of lack of continuing to write on a daily basis has gone to the dogs). With four years in Japanese universities and 20 years living in Japan I had the opportunity to experience what it is to get reasonably proficient in a second language.

The synaptic connections you can make by understanding 'most' of a compound kanji (multiple characters joined together) along with parts of the kanji makeup for a specific kanji you do not know puts you in an incredible position to figure out what the word is likely to mean.

Regardless, once you Japanese gets to a point, you begin to realise it is less about the meaning of words and more about the context and usage of language. You find yourself understanding how to use some words you have not studied without having said to yourself 'so what is this word equivalent in English' - and at times there is not a word equivalent in English. That is when you realise you are thinking in Japanese not just translating.

Synaptic connections, context and knowing what the relevant responses to how you want to respond are the key to fluency in my opinion. This is also the case for writing emails etc.

This also explains why people who do homestays where they get to watch people interact with each other progress faster than someone who goes to live with their significant other half. If you are involved in half of the interactions you get less opportunity to mimic the common responses.

Sorry for digressing a little on this one. Once you get that 777 kanji mark you would be well on your way to being able to contextually understand much more for the kanji/sentences you do not fully comprehend - which in turn leads to more kanji being 'gut-felt' understood if not fully formally learnt.

An additional benefit once you get some proficiency is that you can hear a word you do not know - but you can guess the half of the kanji being used in the compound word (and confirm by asking) and that gives you some context of what the word means. Very much like using latin/greek roots in English to breakdown words.

Time to go back to my coding before the day gets away from me. I just needed to respond to this one because I thought Diego's opinion on language learning was too far off the mark from my experience with it.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#189

Earlier quoted context omitted.

I don't disagree with your point in general but the irony is that your first sentence is an excellent demonstration of the parent comments point. I have no clue what a "N2" is and I'm not entirely sure what you meant by "Kanji/readings" . There are only six other sentences of context and they didn't really help to understand the first one.

And still you're fluent in English (I assume) - it's not an issue with the English language but your understanding of the context. Japanese (and Chinese) work a bit differently than latin languages. Even if you know 100% of the kanji (very few Japanese people do), it doesn't mean you know 100% of the words - and vice versa. Since the characters are idiomatic, it also makes it easier to guess the meaning of a word or…

I think your estimate for N2 is a bit too optimistic: passing it in two years would be very difficult without living in Japan.

For context, it took me 3 years of university (outside Japan) to pass N2.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#190

I know around 10,000 Japanese words. As far as kanji goes, I've studied some 1200 of them intensively; through vocab I know many more as parts of words. I quite often have to reach for a dictionary when reading. If you know 777 kanji in some way (like associating them with meanings, through your native language) and you haven't crammed on any vocabulary, you absolutely will not be able to read a thing. In fact, even…

> I quite often have to reach for a dictionary when reading So when reading kanji, how do you look up a word (picture) you don't know? Since there isn't a minimal set of characters, the notion of "alphabetical order" seems impossible. Weird that I've never thought about this until now, but I'm honestly baffled.

These days, if I'm reading print, I use the kanji appendix of the dictionary I'm using (big red Kokugo Jiten, by Kodan-sha, 1993). The kanji are grouped by stroke count, then within stroke count by radical. Usually kanji lookup can be avoided.

Firstly, unusual words tend to have furigana, which makes it trivial. Next, if the unknown kanji isn't the first one in the word, it's possible to do a prefix-based dictionary lookup. In many cases, also, I've been able to guess a reading by common structure. E.g. both 赤(red) and 跡(traces, remains) have a "seki" reading due to a common element, and 根(root) and 痕 (traces, remains) have a "kon" reading, also due to a common element. You might be able to guess at 痕跡 (konseki) by thinking of 根赤.

How we can find 跡 in the Kokuko Jiten's appendix is by counting the strokes first: 13. Then in the 13 stroke section, of the appendix, we find the subsequence of 13-stroke kanji that have the 足 seven stroke radical. The radicals are sorted by stroke count also, so we can find this subsequence fairly quick.

I can recognize quite a number of words that have at least one kanji which doesn't occur in any other word that I know. I've never studied the kanji in isolation, but I can recognize it in that word.

Post reply on HN