Live data from Hacker News

The most frequent 777 characters give 90% coverage of Kanji in the wild

japanesecomplete.com

11–20 of 214 posts

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#11
post #6

The thing is, 90% coverage is not that great. What happens is that you understand common words that make up for a lot of structure, but when an uncommon word appears it's probably important to the sentence. For example, "son, if you go to the plumbf tomorrow morning don't forget to pick up some zlonks." 98% is closer to what you need in order to read a text and have an idea of what's going on. See this article: https…

This is why picture books are so important for improving listening comprehension for young children. It gives them something to focus on so that hearing words they don’t understand or new grammatical constructions is less distracting, it provides visual cues related to some of those words and helps guide understanding of the story, and it gives the children and whoever is reading to them something to point at in discussion. It also makes the stories more appealing and makes repeat readings more rewarding.

I would guess (based on not-too-rigorous anecdotal experience) that the optimal level of word familiarity for most rapidly improving listening comprehension when read-to 1:1 is closer to 90% than 98% (but with the caveat that books should be re-read several times). The big difference is that having someone to talk to makes it very quick and easy to ask what words mean and discuss other aspects of the story.

When read to for an average of 1+ hour per day, young children improve at listening comprehension very quickly.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#12
post #6

The thing is, 90% coverage is not that great. What happens is that you understand common words that make up for a lot of structure, but when an uncommon word appears it's probably important to the sentence. For example, "son, if you go to the plumbf tomorrow morning don't forget to pick up some zlonks." 98% is closer to what you need in order to read a text and have an idea of what's going on. See this article: https…

As someone who barely scraped by the Kanji/readings of my N2 but have to do a chunk of my work in Japanese, I gotta disagree.

Sure, reading isolated sentences with only 90% coverage is really hard sometimes. But usually you're reading whole blocks of text, so you have context. Also, maybe you don't know some kanji but it has similar radicals to other, so you can make educated guesses (only sometimes of course).

I mean it's just like English! There are radicals/etymological clues, there are context clues, there's so much that you can work off of to figure things out. And exercising "figuring things out with partial information" is a valuable skill that can compensate for a lot of missing kanji knowledge.

An aside: I've heard that when learning a language, you should be able to understand roughly 80% of what you're reading/listening to, because then you'll have enough of a base to not feel overwhelmed, while still having new stuff to absorb.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#13
post #8

I thought the number of standard kanji was only about 1900 to begin with.

2136.

But the list isn’t at all comprehensive. There are a considerable number of kanji in regular use that aren’t on the list, and when you include place/person names, it grows massively.

It’s possible to memorize all the standard kanji, but crack open a history book and you won’t recognize half of the words.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#15
post #5

Zipf’s law

Beat me to it. To expound upon this, Japanese is not unique in basically adhering to Zipf's law. In many organization data sets, including the vocabulary of most languages, the most commonly used word is twice as common as the second-most, and the second-most word is twice as common as the third-most, and so forth.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#16
post #6

The thing is, 90% coverage is not that great. What happens is that you understand common words that make up for a lot of structure, but when an uncommon word appears it's probably important to the sentence. For example, "son, if you go to the plumbf tomorrow morning don't forget to pick up some zlonks." 98% is closer to what you need in order to read a text and have an idea of what's going on. See this article: https…

You say that like it's a bad thing. I wish I could decipher that much of a random sentence in Japanese. At least then I can ask about the parts I don't know. And for most people learning Japanese, they'll probably already know the word in hiragana or katakana, knowing kanji is just icing on the cake.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#17
post #14

Lol, 777 kanji is just too short to even be comfortable reading emails at work in Japanese. So what good is 90 percent if the only thing it allows you is to do shopping?

It gives you 90% coverage in general. That's not 100% for any given domain. You're likely always going to need to know something beyond that common core for any particular domain. However, that 10% is going to be specific to the domain, and not have much overlap with other domains. The 10% that will be helpful for you at work is likely going to be different from the 10% the tI need for work. And you can get a long way on knowing 90% of something, and have the skills to be able to ask questions about the remaining areas of ambiguity.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#18
post #6

The thing is, 90% coverage is not that great. What happens is that you understand common words that make up for a lot of structure, but when an uncommon word appears it's probably important to the sentence. For example, "son, if you go to the plumbf tomorrow morning don't forget to pick up some zlonks." 98% is closer to what you need in order to read a text and have an idea of what's going on. See this article: https…

For Japanese it can be different, as most structural words are in Kana, a phonogram, instead of in Kanji.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#19
post #7

Vocabulary frequency lists have long been very popular with publishers, so I assume as well with language learners. The assumption behind these is that with knowledge of, say, 90% of the vocabulary you encounter, you will comprehend 90% of the writings you encounter. It doesn't work out at all that way, though, because it is generally the low frequency vocabulary that carries the key information in a text and on whic…

In the US, there are beginning reader books (e.g. the “I Can Read!” and “Ready to Read” series) which intentionally use somewhat limited vocabulary. These are somewhere between a picture book and a chapter book: they have usually 1.5 pages of text and 0.5 pages of picture in each 2-page spread; the text is set in a large font but there are at least a few sentences per page; usually the books are 30–50 pages long, with 3–5 “chapters”.

I don’t know how useful these are for independent reading by 5–6-year-olds, but anecdotally they are great material for reading to 2-year-olds, better than most picture books. (Note: some of the recent readers are garbage marketing gimmicks with movie tie-ins, ranging from boring to incomprehensible; skip those.)

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#20
post #5

Zipf’s law

Beat me to it. To expound upon this, Japanese is not unique in basically adhering to Zipf's law. In many organization data sets, including the vocabulary of most languages, the most commonly used word is twice as common as the second-most, and the second-most word is twice as common as the third-most, and so forth.

> second-most word is twice as common as the third-most, and so forth.

Normally Zipf's law refers to the frequency being inversely proportional to rank - i.e. the 3rd most common element would be 1/3 as frequent as the first, not 1/4th.

Post reply on HN