Live data from Hacker News

Sorting in Japanese – An Unsolved Problem (2011)

localizingjapan.com

11–20 of 120 posts

Re: Sorting in Japanese – An Unsolved Problem (2011)

#11
post #3

https://en.m.wikipedia.org/wiki/Four-Corner_Method You can sort Chinese characters (including Kanji but i'm not sure they use the Four Corners Method) by the Four Corners method. Why would you need to sort kanji phonetically in the first place? Do Japanese users actually expect names to be sorted phonetically? English speakers don't expect names to be sorted by IPA so consistency of the sorting scheme should be all t…

I actually made a serious attempt at learning the Four-Corner Method for kanji [0] and it was very frustrating. It would be difficult to determine what parts of the kanji belonged to which corner, and which exact shape they corresponded to. And strokes wouldn't always be interpreted the way I thought they'd be since it's based on a handwritten representation of the character. Many characters also have multiple FC numbers! The FC method was never meant to uniquely identify specific characters, but just to help narrow down a list of candidates in a dictionary. Funnily enough, I also argue something similar in this thread that has the same drawbacks :)

[0]: Because I was interested in typing characters while not actually knowing the kanji. The Tagaini Jisho app (https://www.tagaini.net/) was indispensable because it lets you search on multiple parameters including partial FC # and simpler methods like SKIP codes (http://nihongo.monash.edu/SKIP.html). The only characters I couldn't transcribe with this method were those printed so small that the individual strokes were difficult to make out.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#12

Am I missing something subtle in this Kanji example, or all 4 names actually written the same?: "There are four Japanese women whose names you have to sort: Junko, Atsuko, Kiyoko, and Akiko. This does not seem difficult, until they each show you how they write their names in kanji: 淳子 (Junko) 淳子 (Atsuko) 淳子 (Kiyoko) 淳子 (Akiko)" I'm not familiar with Japanese at all (and have never had to deal with localization beyond…

[deleted]

Re: Sorting in Japanese – An Unsolved Problem (2011)

#13
>In English, the logo goes with their saying: “Everything from A to Z.” This is indicated by the arrow. But in Japan, and any other country that doesn’t use English, A and Z aren’t always the first and last letters of the alphabet.

This is grasping for straws.

a) Many English speakers/countries are unfamiliar with the meaning of the logo's arrow. b) They will get the meaning if explained. c) Some won't because A-Z standing for "everything" requires a given level of literacy (cf. alpha and omega) d) They will get it easily when explained because it's a simple concept.

Japanese people learn the English alphabet pretty early on in school and while some may not be familiar with the saying A-Z the logo and the meaning still perfectly works and would get an aha reaction when explained.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#14
post #3

https://en.m.wikipedia.org/wiki/Four-Corner_Method You can sort Chinese characters (including Kanji but i'm not sure they use the Four Corners Method) by the Four Corners method. Why would you need to sort kanji phonetically in the first place? Do Japanese users actually expect names to be sorted phonetically? English speakers don't expect names to be sorted by IPA so consistency of the sorting scheme should be all t…

I actually made a serious attempt at learning the Four-Corner Method for kanji [0] and it was very frustrating. It would be difficult to determine what parts of the kanji belonged to which corner, and which exact shape they corresponded to. And strokes wouldn't always be interpreted the way I thought they'd be since it's based on a handwritten representation of the character. Many characters also have multiple FC num…

Don't know about Japanese, but other well-known input schemes for Chinese includes Cangjie (https://en.wikipedia.org/wiki/Cangjie_input_method), Zhengma (https://en.wikipedia.org/wiki/Zhengma_method), and Wubi (https://en.wikipedia.org/wiki/Wubi_method).

In fact, all of these non-phonetic input/encoding systems are highly non-intuitive and have a reputation of sharp learning curves, frustration is expected. This is because, in general, Chinese or Kanji characters are expected to be pronounced or written by the speakers, not to be indexed in a particular encoding system. Only the pronunciation is the natural form in the language.

The encoding schemes are completely foreign, arbitrary to the native speakers. Using them requires extensive and systematic training. In mainland China, Hong Kong and Taiwan, in the 80-90s, learning to use a computer often starts from learning the code, and it needs at least two months of mechanical memorization to get started, and years of use to master it, just like how amateur radio operators learn Morse Code (edit: well, you don't need to memorize the code for every single character as if it's Morse Code, but remembering the standard decomposition of characters in the system is comparable to rememebr the Morse Code table). And remember, these are native speakers, much greater effort is needed for foreign speakers.

Sure, schemes based on radicals have been used for a thousand years in dictionaries, but all of these schemes used today are a completely artificial creation for typing and searching things into/from computers (often on those with very limited computing power).

The increase of processing power of personal computers in the late 90s allowed phonetic input systems to map pronunciation to characters heuristically, with high correctness rate. So those codes are rarely used by Chinese, and Japanese speakers (I believe) today.

Unless you've learned computing in the during 80s to mid 90s, or you have a job related to language or word processing that requires typing tens of thousands of characters or creating/searching them in a language-related database, or you are someone who emphasize typing efficiency.

> Because I was interested in typing characters while not actually knowing the kanji.

This is actually a common requirement for people with those jobs, and one of the biggest reason to keep using them. It is especially useful when transcribing texts to computers or searching them in a database.

Users also argue that, using them help preventing the modern disease of forgetting the writing of characters due to computerization, which I do see a point, similar to the spell-checker problem in English education.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#15

Am I missing something subtle in this Kanji example, or all 4 names actually written the same?: "There are four Japanese women whose names you have to sort: Junko, Atsuko, Kiyoko, and Akiko. This does not seem difficult, until they each show you how they write their names in kanji: 淳子 (Junko) 淳子 (Atsuko) 淳子 (Kiyoko) 淳子 (Akiko)" I'm not familiar with Japanese at all (and have never had to deal with localization beyond…

> How do you get further context on how to pronounce a proper name like this?

You don't. There's no way to know which pronunciation to use other than the person telling you.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#16

Am I missing something subtle in this Kanji example, or all 4 names actually written the same?: "There are four Japanese women whose names you have to sort: Junko, Atsuko, Kiyoko, and Akiko. This does not seem difficult, until they each show you how they write their names in kanji: 淳子 (Junko) 淳子 (Atsuko) 淳子 (Kiyoko) 淳子 (Akiko)" I'm not familiar with Japanese at all (and have never had to deal with localization beyond…

On Japanese forms that require you to use your name, you provide them with how your name looks in kanji, as well as how they are read (in their syllabic alphabet). When introducing yourself in speaking, you may also mention how your name is written.

It's not a problem in the sense that when you are in a position to ask someone for their name, you are also in a position to ask them for both the orthographic and phonetic versions of their name; it's just not something speakers most other language are familiar with handling.

Also note that most Japanese people you encounter will have names with a couple obvious readings. In the 淳子 example, Junko should be the most likely reading, followed by Atsuko, seeing as "atsu-" is typically associated with a different kanji. Similarly for Kiyoko and Akiko; in usage out of names, "kiyo-" and "aka"/"aki" are primarily written with other kanji. Since all "native-Japanese" readings for the word are typically written with other kanji, and this kanji is rarely seen in normal text with a native reading, what's left is the most common "Chinese" reading for the word, "jun".

There are more "species" of Japanese personal names, less common, including names written completely in their alphabet, or names where kanji is used only for their phonetic reading so that when written together it forms a native-Japanese word (and hence name), or archaic names with archaic or metaphorical readings, or names of foreign East-Asians that are read as an approximation of how they are pronounced in their respective native language.

Put together, most names have an obvious reading and with enough time in Japanese society you should know the obvious exceptions, but the rules for pronunciation are still sufficiently haphazard and irregular that asking for pronunciation is necessary to be sure. c.f. mispronunciation of Saoirse in English for a related phenomenon. In forms,one still does better by asking for both orthographic/phonetic rather than having a huge rulesets and tables to divine pronunciations from orthography (with necessary mispredictions), except perhaps in interactive contexts like in IMEs.

You may think this all is complex, but when it comes to names, there is always a lot of complexity to it based on historical linguistic and orthographic phenomena, but one is blind to them when one grows up in them. Consider that the pronunciation of many English cities are not what you'd expect from just looking at how they are written, or that York comes from Proto-Celtic Eburos + akom -> Latin Eboracum -> Old English Eoforwic -> Norse Jorvik -> Middle English York. Also consider that in many languages, e.g. Latin, words need to be memorized in more than one form, since historical processes make it that the stems for different tenses cannot be regularly constructed.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#17
post #3

https://en.m.wikipedia.org/wiki/Four-Corner_Method You can sort Chinese characters (including Kanji but i'm not sure they use the Four Corners Method) by the Four Corners method. Why would you need to sort kanji phonetically in the first place? Do Japanese users actually expect names to be sorted phonetically? English speakers don't expect names to be sorted by IPA so consistency of the sorting scheme should be all t…

Imagine having your contact list sorted by some geometric function run on each letter of the A-Z alphabet.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#18

Earlier quoted context omitted.

I actually made a serious attempt at learning the Four-Corner Method for kanji [0] and it was very frustrating. It would be difficult to determine what parts of the kanji belonged to which corner, and which exact shape they corresponded to. And strokes wouldn't always be interpreted the way I thought they'd be since it's based on a handwritten representation of the character. Many characters also have multiple FC num…

Don't know about Japanese, but other well-known input schemes for Chinese includes Cangjie ( https://en.wikipedia.org/wiki/Cangjie_input_method ), Zhengma ( https://en.wikipedia.org/wiki/Zhengma_method ), and Wubi ( https://en.wikipedia.org/wiki/Wubi_method ). In fact, all of these non-phonetic input/encoding systems are highly non-intuitive and have a reputation of sharp learning curves, frustration is expected. Thi…

Phonetic systems for Chinese aren't so prolific for non-mandarin speakers.

9 square/Q9 is another popular method.

The stroke-based methods aren't so arbitrary - they are based on the way you write.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#19

Earlier quoted context omitted.

I actually made a serious attempt at learning the Four-Corner Method for kanji [0] and it was very frustrating. It would be difficult to determine what parts of the kanji belonged to which corner, and which exact shape they corresponded to. And strokes wouldn't always be interpreted the way I thought they'd be since it's based on a handwritten representation of the character. Many characters also have multiple FC num…

Don't know about Japanese, but other well-known input schemes for Chinese includes Cangjie ( https://en.wikipedia.org/wiki/Cangjie_input_method ), Zhengma ( https://en.wikipedia.org/wiki/Zhengma_method ), and Wubi ( https://en.wikipedia.org/wiki/Wubi_method ). In fact, all of these non-phonetic input/encoding systems are highly non-intuitive and have a reputation of sharp learning curves, frustration is expected. Thi…

Based on extremely limited Googling, one of the cases where these codes are still used is written colloquial Cantonese, which lacks any major official support.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#20
post #3

https://en.m.wikipedia.org/wiki/Four-Corner_Method You can sort Chinese characters (including Kanji but i'm not sure they use the Four Corners Method) by the Four Corners method. Why would you need to sort kanji phonetically in the first place? Do Japanese users actually expect names to be sorted phonetically? English speakers don't expect names to be sorted by IPA so consistency of the sorting scheme should be all t…

> Do Japanese users actually expect names to be sorted phonetically?

Yes - or that's how a human would sort them, at any rate.

Post reply on HN