Live data from Hacker News

The most frequent 777 characters give 90% coverage of Kanji in the wild

japanesecomplete.com

201–210 of 214 posts

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#201

Earlier quoted context omitted.

HN is full of the dilettante type who consider themselves an authority on any field based on a few days of wiki-hopping. I notice this every time I see comments on HN about a topic I actually understand in depth. Quite a few comments are just outright incorrect, most are misguided, but all are confident.

Just because someone posts a comment with their understanding of an issue doesn't mean they consider themselves an authority.

Yes, but a lot of people who are aware of their lack of understanding (or reduced understanding) actually choose to stay silent, or offer opinions with a lot of riders.

On the other hand, the most confident pronouncements usually come from people unaware of their lack of understanding.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#202
post #116

In other news only 24 english characters give nearly 100% of all the characters in use in English

But knowing 24 characters doesn't make you understand 100% of the words. Actually, it gives you 0%.

That was the point.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#203

Earlier quoted context omitted.

How do you check your coverage? Or even how many words you know?

Just googling around, I found this test: https://www.arealme.com/vocabulary-size-test/en/ No idea about its accuracy, but as a native English speaker it's telling me my vocabulary size is 30k words, which sounds roughly correct.

I've tried it now and got 22k, which seems not so bad for a foreigner ("Top 6.53% Your vocabulary is at the level of professional white-collars in the US!"), but I feel like I cheated: most of the more fancy English words are just misspelled Latin, and having even a modest Latin vocabulary (I'd don't think I know more than 4k Latin words) makes their meanings pretty obvious.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#206

Earlier quoted context omitted.

> the characters got arbitrarily attached to an existing word because they were similar in sound and meaning

It does not seem arbitrary in this case because the meaning does match. I do take your point that using that word in the discussion above, which is about Japanese was not the best example. On the other hand, it is a good example in chinese.

(A) In Japanese the meanings don't match that closely. The word didn't originally mean cute, but rather pathetic or pitiable, and evolved over time. More info: http://gogen-allguide.com/ka/kawaii.html

(B) By arbitrary here I mean that there is no linguistic connection. "Arbitrarily chosen because they are similar" => "chosen for no reason other than their similarity".

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#207
post #198

Before this falls off the front page, I figured I would ask the following: 1) Does anyone have a resource for Japanese subtitles (in Japanese/kanji, not English)? 2) Does anyone have good frequency { word => frequency } lists? Especially if they are topical, eg. school-related, anime-related, industry-related. 3) What are the best programs for segmenting Japanese text into words reliably? 4) Does anyone have a vocabu…

5) zkanji: https://github.com/z1dev/zkanji ! It's got a dictionary from which you can directly add words into its study decks when looking them up, and it has handwriting recognition plus let's you easily find similar looking kanji (with shared components etc), which is great when tesseract ocr fails, or when the text is so blurry/compressed you can't really even see it clearly (the online Japanese war history archiv…

Thanks for the head's up! :)

Not long after this HN thread, this thread popped up on Reddit. It answered a lot of my questions ([1], [2], [4], and [7]), and I found it immensely useful:

https://www.reddit.com/r/LearnJapanese/comments/crlsqj/googl...

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#208

Earlier quoted context omitted.

And still you're fluent in English (I assume) - it's not an issue with the English language but your understanding of the context. Japanese (and Chinese) work a bit differently than latin languages. Even if you know 100% of the kanji (very few Japanese people do), it doesn't mean you know 100% of the words - and vice versa. Since the characters are idiomatic, it also makes it easier to guess the meaning of a word or…

> "N2 is generally attained after 1-2 years of Japanese studies" From my own experience, this would only be true in the most favorable of the situations. 2 years studying fulltime while living in Japan sounds about right. 1 year studying as a hobby few hours a week, no way.

Oh yeah, I definitely meant full-time studies coupled with practice outside of study-hours.

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#209

Earlier quoted context omitted.

That's if you know 90% of the words. If you recognize 90% of the kanji and know what they mean, that's not the same as knowing words.

> That's if you know 90% of the words No, I'm taking about the premise of the article: if you know 90% of the words most widely used (that is, the above 777 characters for example). Not 90% of the words in general. If you know the most likely words, you can more often than not guess the unknown words in a phrase from the context.

[deleted]

Re: The most frequent 777 characters give 90% coverage of Kanji in the wild

#210

Earlier quoted context omitted.

That's if you know 90% of the words. If you recognize 90% of the kanji and know what they mean, that's not the same as knowing words.

> That's if you know 90% of the words No, I'm taking about the premise of the article: if you know 90% of the words most widely used (that is, the above 777 characters for example). Not 90% of the words in general. If you know the most likely words, you can more often than not guess the unknown words in a phrase from the context.

I'm also talking about recognizing 90% in whatever text you're reading, not about knowing 90% of the lexicon!

My point is that those characters aren't words.

Well, some of them are, sometimes.

Post reply on HN