Live data from Hacker News

Sorting in Japanese – An Unsolved Problem (2011)

localizingjapan.com

81–90 of 120 posts

Re: Sorting in Japanese – An Unsolved Problem (2011)

#81
post #76
post #61

Earlier quoted context omitted.

Arguably, emojis are going in that direction. And, seriously, the rest of the world considers it totally illogical to use MM-DD-YYYY. In almost any country, MM are the middle digits. 4月12日2019年 is a format I have never seen used.

Fun fact, emoji was invented by the Japanese. :-)

I wonder what people think its root is.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#82
post #69

As a Japanese native, it pains me a lot when the westerners come upon linguistic differences like this and start calling them things like "overly odd" or "inferior" as they seem to be doing in the comments here. I can't help but smell a whiff of anglophonic condescension - I had thought political correctness have swept over the west over the last few years, but perhaps they only selectively apply to those who can tal…

Reading the article showed me how ignorant I was to this problem... so much so that I'm surprised I haven't seen more articles like this in the past. For all the talk on ML and PC here, it's surprising that language is still a huge unsolved issue. As someoneone who knows a tiny amount of Japanese, I'm surprised that a common solution is to provide the katakana equivalent. Outside of the tech realm, is this really wha…

Yes, entering a name's kanji and kana separately is standard for most any form in Japan, whether digital or on paper.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#83

Earlier quoted context omitted.

And as a counter example, I like kanji and learned Japanese by reading (though I also like manga, so perhaps it was helpful). I find kanji very helpful in learning vocabulary. In fact, I actually measured how fast I could learn vocabulary containing kanji I didn't know by learning the kanji first, versus learning the vocabulary phonetically. It was about the same. However, as there are only about 2200 kanji in the li…

I learned just over 17,000 words and for that needed to know somewhere north of 3000 kanji. ~2200 is just the joyou list but once you get into novels and the like you can expect to encounter i'd say around 3500 would be close to maxing it out with the last 0.01% being another couple of thousand, but that's getting in to really obscure shit that only a smattering of the general population know. I learned kanji first b…

It's been years since I did any studying. You've inspired me to start again :-)

Re: Sorting in Japanese – An Unsolved Problem (2011)

#84
Some statements you have made about the Japanese Language are ambiguous at best. Let's clear them up, but first let's be precise with the terminology. Pronunciation and sound have not interchangeable meaning. Every sound -or sequence of- in Japanese is associated with a specific character. It happens that some Kanji have multiple sounds or that a group of Kanji shares the same sounds, but still how are read it's univocal. There is no pronunciation involved. Just to give a quick example, the sound of あ is one and only one, while in English the sound of `a` has different pronunciations in pal, Paul, paediatrics, and so on. For the sake of simplicity it's ubiquitous savage practice overlapping the use of the "reading" with the one of "pronunciation" and the one of "intonation", but you should know the difference. The pronunciation transforms based on the preceding and/or following characters. There's no such concept in Japanese. The sounds of the Japanese language are distinctively dictated by Hiragana. Therefore what you have is a "Reading". A reading is made of one or more sounds. But once again readings or sounds have no multiple pronunciations. I hope I was able to shed some light on this complicated subject. It's a nuance, but makes an important impact.

Quoting: "I should note that there are two different alphabetical sorting orders in Japanese. For this article I am going to use the a i u e o (あいうえお) sort order." Alphabetical order as we know it, it's one in Japan too. The order of the Hiragana works exactly like our alphabetical order: a preordered sequence of scripts from A to Z is the methodologically equivalent to the Hiragana from あ to ん, albeit there is no words starting with を and ん - last two characters of the alphabet. The Katakana alphabet is exactly like Hiragana, just the characters `design` changes - it is only used to write, in Japanese sounds, imported words from non-Japanese languages (Chinese and Korean are exception because you can use the Kanji to write them, although when writing just Chinese or Korean sounds, you will be using Katakana still). When talking about `ordering` we normally refer to which sequence we are listing the Kanji: by their Hiragana sounds is the most common way of listing them. They can also be listed by what in English are called Radicals - foundational shapes composing the Kanji ideograph, they can be listed by strokes count or be listed by their Chinese reading - the Japanese sound of the equivalent Chinese Kanji, or even by recurrence. In the Japanese language it's frequently used `ordering` by meaning, by which Kanji may or may not be listing, where only Hiragana is used to write all the words. This method is the closest rendition in Japanese to what we are accustomed to name English dictionary. More here https://en.wikipedia.org/wiki/Japanese_dictionary I think it's already clear where the problem of sorting lays. Kanji, Hiragana and Katakana, though being different looking alphabets, have a well-defined scope and interconnection in the Japanese language. Kanji are words, thus have meaning; Hiragana offers a vocalization to the Kanji and also absolves the crucial grammatical role, providing the language with adverbs, prepositions, particles, determiner, verbs conjugation and more. Katakana usage is sidelined to just words foreign in nature, hence they exclusively represents sounds, carrying no meaning. Consequently when reading Japanese on a generic subject Kanji are mostly encountered, some Hiragana that connects them are present (it needs to be said that very common words are often written in Hiragana only, e.g. こんにちは means Hello) and sporadic Katakana when some imported word is used. This is happening because in the Japanese language there is hardly any punctuation at all - yes, there's a full stop, a comma, a way to encapsulate direct speech, but no concept of empty space. You can read Japanese from any direction you set up yourself for. Traditionally it's read top to bottom from left to right. Modern books are read as ours would be, even so sometimes you turn the pages backward to further the reading, e.g. comics etc. The English alphabet is used in case of very technical terms or for proper names - the context may sometime better served with romaji. But that's not a rule whatsoever. Katakana is employed more often. The Western alphabet -romaji in Japanese- adoption is similar to how you would take up French, Spanish, Italian or German words in your writing. (continue)

Re: Sorting in Japanese – An Unsolved Problem (2011)

#85

Some statements you have made about the Japanese Language are ambiguous at best. Let's clear them up, but first let's be precise with the terminology. Pronunciation and sound have not interchangeable meaning. Every sound -or sequence of- in Japanese is associated with a specific character. It happens that some Kanji have multiple sounds or that a group of Kanji shares the same sounds, but still how are read it's univ…

Quoting from "Sorting Settings": "In this example you can see ABC and katakana are separated. Kanji are then separated from katakana. There were no hiragana in this list[...]" The Hiragana on that list are the characters と and の. On that particular list の expresses the meaning of correlation and と convey the mere meaning of `and`. For instance 地域 と 言語 の オプション (notice I myself have entered the `spaces` for clarity's sake, but there should be none) is a very instructive example. Those are actually three distinct nouns you are trying to order as one: 地域 reads ちいき (Hiragana for `chiiki`), means area; 言語 reads げんご (Hiragana for `gengo`), means language; オプション read opushon, is, you guessed, option in English. Note I didn't write that オプション means option. オプション IS the word `option` written with Japanese sounds or, in other words, written in the Japanese script - i.e. Katakana, you guess it. So 地域と言語のオプション can be translated as "Regional and Language options".

Quoting from "Sorting Names" "It is very possible to have different people with the same name write their name in different character sets. The traditional way of writing the Japanese name of Ayumi would be written in kanji; a modern, stylish way would be to write it in hiragana, and a second generation Japanese-American might write their name in katakana or the alphabet." Japanese always writes their name in Kanji. They don't use different sets. When the Kanji composing their names can assume several sounds, they write alongside what's called Furigana, the Kanji reading - usually as a subscript or superscript. Furigana is written in Hiragana (not Katakana as you stated later on), but for the Internet where websites could potentially be read by a non-Japanese crowd, Katakana might be used in some cases. Nonetheless as I said earlier Hiragana and Katakana differ only by scripting style, if you wish to call it as such. So which script is in use, it's not a relevant issue for foreigners. All the other statements are a matter of opinion, save for the last: Japanese do write sometimes their names in the romaji, aka "ABC alphabet", especially when they are dealing with foreigners at any level. Although why, being 2nd gen Japanese-American entails writing its own name in Katakana, beats me. It's a tiny bit like saying, forgive me here, a 2nd gen Italian-American would write its name in Latin.

Quoting from "Kanji - The Real Problem": "Kanji have multiple pronunciations, determined by the context in which it appears.[...]Only from the context in which the kanji appears do you know how to pronounce it." That's like saying the pronunciation of the word 'pool' is determined by the context referring to water or balls. Not the pronunciation, but the meaning of the word is changed by the context. We all agree on this statement. Single Kanji words change along with their meaning based on the context, but like the word 'pool', its reading is the same. e.g. あめ (reads ame) which means 'candy' if written 飴 or 'rain' if written 雨 . In this case the Kanji itself controls its meaning and reading, not the context. When single Kanji is a verb the meaning and reading changes with what's called Okurigana - Hiragana written after the Kanji absolving the purpose of conjugation -, while the Kanji remains invariant. For example 着く and 着る reads つく (tsuku) and きる (kiru) respectively, the former means 'to arrive' and the latter 'to wear'. Single Kanji aren't at all like you describe them. Compound Kanji words follow a different ruleset build upon how many they are. Generally speaking one kanji in two-Kanji words has multiple readings depending on what is the word it appears in and where it appears in that word. You can learn the rules, or you can get used to them just by seeing them used in massive frequency. This case alone is as you've described: the need to know the context in which the Kanji lives. But since the meaning also changes with its reading, you could be able to catch the overall meaning of the sentence without being able to read that single word. But compounds Kanji with multiple readings aren't that recurring and they generally represents common words requiring not much effort to memorize. Compound Kanji words composed by more than two Kanji reads unequivocally in one way as English words do with very few exceptions. More here https://ja.wikipedia.org/wiki/%E5%90%8C%E5%BD%A2%E7%95%B0%E9... If anything what could, quote "[...]keeps students up nights studying for years[...]" isn't the multiple Kanji readings, but the fact that you need to know between 2000 and 3000 Kanji and its their combinations that build words (mostly two Kanji words). So it's like having a permutation with repetition (in ordered arrangements) of 3000 syllables that makes words in pairs or singularly.

Quoting "Here is an example: 私は私立大学で勉強しています。[...]A second year Japanese student could figure this out. For a computer, this is a very difficult problem." The choice is particularly sad. This isn't difficult at all for a computer, granting you understand the Japanese language. Let's dissect this phrase: 私 は 私立大学 で 勉強しています。 私 わたし (reads watashi) is the English pronoun 'I'. A computer instantly knows it because only the watashi reading/meaning can be followed by the Hiragana は. That's something a 6 years old Japanese knows. And something you would learn in your first months or less of Japanese studies. Conversely a computer know instantly that's the reading/meaning isn't watashi when it scans that following 私, there's another Kanji. This compound Kanji 私立 reading can only be しりつ (reads shiritsu) out of a staggering number of combined readings of 2 - and that's only because 立 has two usable Chinese reading りゅう and りつ (the third would require the Kanji to be lonesome), 私 only one, し. Kanji have usually one Japanese reading and one or more Chinese readings governed by strict rules on which reading group has to be used. Coding it isn't as much of a headache. As soon as the computer realize that the third character that follows the first two Kanji is a Kanji as well, the range of possible readings bottoms. That's also due to the fact the first two Kanji makes already a word - as often happens with 2+ long Kanji words, they are compose of multiple words, just like some long words in Western languages would - that means 'private' in English. With the same approach the computer instantly finds the reading of the two Kanji 大学 だいがく (reads daigaku) means University, another very common noun. I think you already got the gist of it. Last word 勉強しています the computer know instantly is a verb because of the unique okurigana しています (read shiteimasu) Present Continuous of "to do" and 勉強 is both extremely popular and has a unique reading, べんきょう (reads benkyou) which is the noun 'study'. "I'm studying at a private university", even a machine translation would be accurate here.

I think the point is that there is no use in sorting all the words written in the three different Japanese alphabets simultaneously in the same juncture. Microsoft knows it so well it has yet to implement it. In your final thought you completely miss to understand that you don't need to attack the problem by "pronunciations". You have only to treat the Kanji with different approach and translate Hiragana/Katakana in romaji, which it has been done already long time ago. I hope at least you're going to quit using "pronunciation" in favor of "reading" by the time you've done reading this post. If ever.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#87
post #69

As a Japanese native, it pains me a lot when the westerners come upon linguistic differences like this and start calling them things like "overly odd" or "inferior" as they seem to be doing in the comments here. I can't help but smell a whiff of anglophonic condescension - I had thought political correctness have swept over the west over the last few years, but perhaps they only selectively apply to those who can tal…

I don't think languages deserve any special protection. People do.

Also, I think that all languages are supper weird, and because of that they have their own unique cultural value that can't be measured. But in terms of economical value:

1. English happens to have the most speakers if you sum L1 and L2 speakers (1.132 billion)

2. English has by far the biggest non native speaker representation (753.3 million)

3. Speaking English gives a lot of economical opportunities, because English is widely spoken among the richest people in the world. By rich people I mean top 1% with more than 30 000$ income per year i.e. people living in USA, Western Europe, Australia etc.

Because of these 3 points I do consider all other languages, including my native Polish, inferior as a tools. And I say this even tough I'm fully aware that millions were dying across world and history in defense of their native speech. History of my own nation and my native language is a perfect example. Millions of Poles died in defense of Polish language against Germans and Russians.

"We won't forsake the land we came from, We won't let our speech be buried. We are the Polish nation, the Polish people, From the royal line of Piast. We won't let the enemy oppress us. (...) The German won't spit in our face, Nor Germanise our children, Our host will arise in arms, Spirit will lead the way. We will go when the golden horn sounds. (...)"

But I think it is exactly the fact that we were not able to communicate with each other that lead to the World Wars in the first place.

I'm studding now in English in a super diverse Dutch university with Germans, Russians, Chinese, Koreans, Arabs, Belgians, French and many other nations, and thought that some 100 years ago we would be all trying to kill each other while living in trenches is just unimaginable. English facilitates that and I would want more people across the world to participate. Being defensive and overly proud of our native language will not help with that.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#88
post #87
post #69

As a Japanese native, it pains me a lot when the westerners come upon linguistic differences like this and start calling them things like "overly odd" or "inferior" as they seem to be doing in the comments here. I can't help but smell a whiff of anglophonic condescension - I had thought political correctness have swept over the west over the last few years, but perhaps they only selectively apply to those who can tal…

I don't think languages deserve any special protection. People do. Also, I think that all languages are supper weird, and because of that they have their own unique cultural value that can't be measured. But in terms of economical value: 1. English happens to have the most speakers if you sum L1 and L2 speakers (1.132 billion) 2. English has by far the biggest non native speaker representation (753.3 million) 3. Spea…

Couldn't disagree more. As you said, people do deserve protection, but who do you think that you are protecting when you protect languages? Exactly, when you protect a language you are protecting people who speak it. If you want to get rid of your native language, fine, but do not force other people to do so.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#89
post #37

Blog author here. Surprised to see one of my old posts on the front page while browsing Hacker News. It's interesting to reflect on what has improved since I wrote it, and what has not. Both Android and iOS, for instance, provide mechanisms to get this right, if you know to use them and expose them for those locales (and only those locales). For example, both have a Contact object that contain corresponding phonetic-…

[deleted]

Re: Sorting in Japanese – An Unsolved Problem (2011)

#90

Earlier quoted context omitted.

All of the examples you mentioned are pronounced differently, so I'm not sure what point you're trying to make.

n'a (んあ) and na (な) are pronounced the same (not 1:1) k-ee-y-oo and k-y-oo are distinct pronunciations that are both written as kiyu (きゅ) (also not 1:1, though this could be considered a accent maybe?) I wasn't trying to make a point, beyond "spelling a [any] word is not just the same thing as saying it".

This is coming more from having looked at Japanese in my phonolygy class than learning Japanese itself (although I have done the latter and noticed the distinction there as well) んあ is 2 syllables, な is one. This changes the prounounciation (most notably the length of the nasal stop.

Ki-yu (きゆ) and Kyu (きゅ) arr distinguised by the size of the ゆ character

Post reply on HN