Live data from Hacker News

The most frequent 777 characters give 90% coverage of Kanji in the wild

japanesecomplete.com

141–150 of 214 posts

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#141

Earlier quoted context omitted.

Afterthought: Kawaii (cute) is actually 可愛い, also containing "可", which literally may mean something like "a thing that can be easily loved", or simpler, cute. But you wouldn't be able to guess that if you just know the Kanji.

I don't know for Japanese as meaning sometimes shifts from Chinese, but in Chinese the standard definition of 可 is "can, may, be able to". You obviously learn it by itself but as Chinese words are mostly a combination of 2 characters, you immediately also have to learn e.g. 可以 (can, may, be able to), 可能 (maybe), 可爱 (cute) etc. So someone who's learning characters in order to get 90% coverage (or whatever) would not s…

> the meaning of 可爱 would be fairly straightforward to guess

To be honest this whole thread about 可愛い is more or less bonkers, because it's an ateji. The word's meaning doesn't derive from the characters, the characters got arbitrarily attached to an existing word because they were similar in sound and meaning.

As such, the whole thing is about as meaningful as talking about how easy it is to guess that 珈琲 means "coffee"...

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#142

Earlier quoted context omitted.

I don't disagree with your point in general but the irony is that your first sentence is an excellent demonstration of the parent comments point. I have no clue what a "N2" is and I'm not entirely sure what you meant by "Kanji/readings" . There are only six other sentences of context and they didn't really help to understand the first one.

And still you're fluent in English (I assume) - it's not an issue with the English language but your understanding of the context. Japanese (and Chinese) work a bit differently than latin languages. Even if you know 100% of the kanji (very few Japanese people do), it doesn't mean you know 100% of the words - and vice versa. Since the characters are idiomatic, it also makes it easier to guess the meaning of a word or…

> "N2 is generally attained after 1-2 years of Japanese studies"

From my own experience, this would only be true in the most favorable of the situations. 2 years studying fulltime while living in Japan sounds about right. 1 year studying as a hobby few hours a week, no way.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#143
post #116

In other news only 24 english characters give nearly 100% of all the characters in use in English

But knowing 24 characters doesn't make you understand 100% of the words. Actually, it gives you 0%.

It turns out there are more than 777 words commonly used. Kanji is not different in this respect.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#144
post #12

Earlier quoted context omitted.

As someone who barely scraped by the Kanji/readings of my N2 but have to do a chunk of my work in Japanese, I gotta disagree. Sure, reading isolated sentences with only 90% coverage is really hard sometimes. But usually you're reading whole blocks of text, so you have context. Also, maybe you don't know some kanji but it has similar radicals to other, so you can make educated guesses (only sometimes of course). I mea…

When I was learning English, the most memorable/epic lesson of advanced English I ever took was trying to understand what we were told was an "advanced" poem; the Jabberwock! We were a class of 6-8 students, in pairs making our guesses of what the words that we didn't "understand" meant. We all got pretty advanced ideas and all got to similar conclusions. The teacher would tell us afterwards how those were made up wo…

GEB by Douglas Hofstadter has an interesting section on translations of the Jabberwocky http://www76.pair.com/keithlim/jabberwocky/poem/hofstadter.h...

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#145
post #116

In other news only 24 english characters give nearly 100% of all the characters in use in English

But knowing 24 characters doesn't make you understand 100% of the words. Actually, it gives you 0%.

On the other hand, if you don't know those 24 characters you're going to struggle understanding anything. It's not sufficient but it is necessary.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#146

Earlier quoted context omitted.

I don't know for Japanese as meaning sometimes shifts from Chinese, but in Chinese the standard definition of 可 is "can, may, be able to". You obviously learn it by itself but as Chinese words are mostly a combination of 2 characters, you immediately also have to learn e.g. 可以 (can, may, be able to), 可能 (maybe), 可爱 (cute) etc. So someone who's learning characters in order to get 90% coverage (or whatever) would not s…

> the meaning of 可爱 would be fairly straightforward to guess To be honest this whole thread about 可愛い is more or less bonkers, because it's an ateji. The word's meaning doesn't derive from the characters, the characters got arbitrarily attached to an existing word because they were similar in sound and meaning. As such, the whole thing is about as meaningful as talking about how easy it is to guess that 珈琲 means "cof…

I must say I don't know much about ateji in Japanese.

In this case, though it does seem that the characters where chosen at least partly because of their actual meaning.

It seems that it is both an ateji and a jukujikun [1] because the word does not come from the characters but the characters do have the correct meaning.

[1]https://en.wiktionary.org/wiki/%E5%8F%AF%E6%84%9B%E3%81%84#J...

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#148

This means very little though. Take me as English learner for example. I would say I was only able to understand everyday English without too much of a hassle, after I acquired like around 10k words, which as I just checked had a coverage about 98%+. Noted, it is still NOT enough, actually far from enough. Right now I believe I master around 15k to 20k words, by various estimates, and navigating English on the intern…

How do you check your coverage? Or even how many words you know?

Just googling around, I found this test: https://www.arealme.com/vocabulary-size-test/en/

No idea about its accuracy, but as a native English speaker it's telling me my vocabulary size is 30k words, which sounds roughly correct.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#149
A few years ago I downloaded several hundreds of megabytes of Japanese subtitles, split into 3 categories: live action/drama, anime and foreign film/tv

I’ve listed them in a google sheets together with a few other corpora

https://docs.google.com/spreadsheets/d/1yb5dq4ahdwc_g0aQTL3Y...

Choose the jimaku tab for subtitles to see how big the variation between corpus can be.

According to other comments here, it appears that OP list is based on a newspaper corpus from 1993.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#150
post #52

Disclaimer: native Chinese speaker, knows some Japanese, English sufferer Putting aside the argument of whether removing all Hanzi from Japanese text would actually be more efficient or not, the question to me is: why stop at Hanzi? Why not romanizating all the Japanese literature? Surely almost all the reasoning in favor of getting rid of Hanzi can also apply here? edit: grammar

I'm pretty sure this (or at least sticking to kana) was attempted shortly after 1945 but was abandoned because of the sheer number of homophones.
Post reply on HN