Live data from Hacker News

The most frequent 777 characters give 90% coverage of Kanji in the wild

japanesecomplete.com

51–60 of 214 posts

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#51

I know around 10,000 Japanese words. As far as kanji goes, I've studied some 1200 of them intensively; through vocab I know many more as parts of words. I quite often have to reach for a dictionary when reading. If you know 777 kanji in some way (like associating them with meanings, through your native language) and you haven't crammed on any vocabulary, you absolutely will not be able to read a thing. In fact, even…

> I quite often have to reach for a dictionary when reading So when reading kanji, how do you look up a word (picture) you don't know? Since there isn't a minimal set of characters, the notion of "alphabetical order" seems impossible. Weird that I've never thought about this until now, but I'm honestly baffled.

There are many ways. You can look up by radicals, 4-corner code or simply draw the kanji on the screen.

Better to have instant dictionary lookup if you're on the phone or laptop than a regular paper dictionary.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#52
Disclaimer: native Chinese speaker, knows some Japanese, English sufferer

Putting aside the argument of whether removing all Hanzi from Japanese text would actually be more efficient or not, the question to me is: why stop at Hanzi? Why not romanizating all the Japanese literature? Surely almost all the reasoning in favor of getting rid of Hanzi can also apply here?

edit: grammar

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#53

I know around 10,000 Japanese words. As far as kanji goes, I've studied some 1200 of them intensively; through vocab I know many more as parts of words. I quite often have to reach for a dictionary when reading. If you know 777 kanji in some way (like associating them with meanings, through your native language) and you haven't crammed on any vocabulary, you absolutely will not be able to read a thing. In fact, even…

> I quite often have to reach for a dictionary when reading So when reading kanji, how do you look up a word (picture) you don't know? Since there isn't a minimal set of characters, the notion of "alphabetical order" seems impossible. Weird that I've never thought about this until now, but I'm honestly baffled.

The "alphabetical" order is based on the sound of the word. If you are trying to look up a kanji you don't know in a traditional dictionary, you generally look it up by the radicals (parts) of the kanji or the number of strokes of the character, but most people these days use an electronic dictionary where you draw the kanji to look it up.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#54
I've read through the linked paper and I can't understand where they get the assertion of 777 characters give 90% coverage. The original paper isn't even about that topic, but rather comparing and contrasting a corpus created in 1994 with a corpus created in 1962 and 1976.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#55
post #6

The thing is, 90% coverage is not that great. What happens is that you understand common words that make up for a lot of structure, but when an uncommon word appears it's probably important to the sentence. For example, "son, if you go to the plumbf tomorrow morning don't forget to pick up some zlonks." 98% is closer to what you need in order to read a text and have an idea of what's going on. See this article: https…

Being fluent in Japanese as a second language, I agree with this. It's Zipf's law and sounds great, but 90% isn't as useful as it sounds. You'll mostly recognize a single common Kanji of compound words consisting of 2-3 characters, or common structural words. It's a far cry from being able to understand content you find in the wild. Also, "understanding" a Kanji is an ill-defined term. Most Kanji have multiple meanin…

Afterthought: Kawaii (cute) is actually 可愛い, also containing "可", which literally may mean something like "a thing that can be easily loved", or simpler, cute. But you wouldn't be able to guess that if you just know the Kanji.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#56
post #6

The thing is, 90% coverage is not that great. What happens is that you understand common words that make up for a lot of structure, but when an uncommon word appears it's probably important to the sentence. For example, "son, if you go to the plumbf tomorrow morning don't forget to pick up some zlonks." 98% is closer to what you need in order to read a text and have an idea of what's going on. See this article: https…

especially since it's the highest entropy words that are missing, by definition... so you always get hit the hardest

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#57
post #12
post #6

The thing is, 90% coverage is not that great. What happens is that you understand common words that make up for a lot of structure, but when an uncommon word appears it's probably important to the sentence. For example, "son, if you go to the plumbf tomorrow morning don't forget to pick up some zlonks." 98% is closer to what you need in order to read a text and have an idea of what's going on. See this article: https…

As someone who barely scraped by the Kanji/readings of my N2 but have to do a chunk of my work in Japanese, I gotta disagree. Sure, reading isolated sentences with only 90% coverage is really hard sometimes. But usually you're reading whole blocks of text, so you have context. Also, maybe you don't know some kanji but it has similar radicals to other, so you can make educated guesses (only sometimes of course). I mea…

I don't disagree with your point in general but the irony is that your first sentence is an excellent demonstration of the parent comments point.

I have no clue what a "N2" is and I'm not entirely sure what you meant by "Kanji/readings". There are only six other sentences of context and they didn't really help to understand the first one.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#58
post #46

Earlier quoted context omitted.

The English language is also quite irregular both in spelling and grammer[sic]. Maybe we should start with it given that children around the world "spend countless hours learning this complex (and often irrational) writing system". Disclaimer: native Chinese speaker, non-fluent Japanese speaker, English sufferer edit: disclaimer and grammar...

I'm a native Japanese speaker but as far as the writing system goes, I fully agree with the linguist Geoffrey Pullum's following statement about Chinese characters / kanji: In consequence, this horror-show of a writing system, with its crippling memorization burden for students and malign impediment to progress in science and industry, is the focus of so much intellectual investment and cultural pride that getting ri…

> with its crippling memorization burden for students

I always wonder about something that might sound anecdotal: why isn't there any equivalence of "spelling bee" in Hanzi? Or maybe there is?

> malign impediment to progress in science and industry

What? Do we have any evidence for this?

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#59
post #52

Disclaimer: native Chinese speaker, knows some Japanese, English sufferer Putting aside the argument of whether removing all Hanzi from Japanese text would actually be more efficient or not, the question to me is: why stop at Hanzi? Why not romanizating all the Japanese literature? Surely almost all the reasoning in favor of getting rid of Hanzi can also apply here? edit: grammar

That's essentially what Korean did. They replaced everything with their own morphophonemic orthography.

But even then they still use Hanja to disambiguate sometimes.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#60
post #12

Earlier quoted context omitted.

As someone who barely scraped by the Kanji/readings of my N2 but have to do a chunk of my work in Japanese, I gotta disagree. Sure, reading isolated sentences with only 90% coverage is really hard sometimes. But usually you're reading whole blocks of text, so you have context. Also, maybe you don't know some kanji but it has similar radicals to other, so you can make educated guesses (only sometimes of course). I mea…

I don't disagree with your point in general but the irony is that your first sentence is an excellent demonstration of the parent comments point. I have no clue what a "N2" is and I'm not entirely sure what you meant by "Kanji/readings" . There are only six other sentences of context and they didn't really help to understand the first one.

I don't know what a N2 is either, but it seems obvious that it is some kind of course for learning Japanese or Kanji, and that he did poorly in it. That's good enough for me.
Post reply on HN