Live data from Hacker News

Sorting in Japanese – An Unsolved Problem (2011)

localizingjapan.com

111–120 of 120 posts

Re: Sorting in Japanese – An Unsolved Problem (2011)

#112
post #73
post #59

Earlier quoted context omitted.

You know, I started by loving kanjis too. I like the way they combine, the stories they tell. And I can imagine that when becoming proficient with them, they form a nice cozy system. But face it: you say I have years to learn them. I could learn 3 languages with the effort wasted to learn kanjis (which are half a language because then you also need to learn the pronunciation of words). No, I'll perfect my English ins…

I'm not sure it's even possible to get proficient at spoken Japanese without learning kanji, past some intermediate point. You either pick them up without intending to (enough to read anyway, if not write), or you don't get proficient. The issue being that proficiency in Japanese means navigating endless homophones and kanji compounds, and kanji are what disambiguates that process. Japanese vocabulary has a huge vari…

A typical 12 year old Japanese kid will be fluent in spoken Japanese and pretty bad at kanjis.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#113
post #112
post #73

Earlier quoted context omitted.

I'm not sure it's even possible to get proficient at spoken Japanese without learning kanji, past some intermediate point. You either pick them up without intending to (enough to read anyway, if not write), or you don't get proficient. The issue being that proficiency in Japanese means navigating endless homophones and kanji compounds, and kanji are what disambiguates that process. Japanese vocabulary has a huge vari…

A typical 12 year old Japanese kid will be fluent in spoken Japanese and pretty bad at kanjis.

Standard school curriculum for a 12 year old covers 1000-1400 kanji, and I said "proficiently", not "fluently".

Re: Sorting in Japanese – An Unsolved Problem (2011)

#114
post #97

For those who want to know more about this: The article touches on just the tip of the iceberg. You might think that all you need to do is add an extra field for phonetic readings, and then simply sort on that field, but there are a lot of things that can go wrong. A naive sort (i.e. based simply on character code) will hit the following snags: 1) Hiragana vs Katakana The article focuses more on Kanji vs Kana, but Ja…

I'd assume you'd want ゅ andっ to be added to the kana they're attached to, and then sorted. I'd want my き and きゅ together and び and っび together.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#115
post #88
post #87

Earlier quoted context omitted.

I don't think languages deserve any special protection. People do. Also, I think that all languages are supper weird, and because of that they have their own unique cultural value that can't be measured. But in terms of economical value: 1. English happens to have the most speakers if you sum L1 and L2 speakers (1.132 billion) 2. English has by far the biggest non native speaker representation (753.3 million) 3. Spea…

Couldn't disagree more. As you said, people do deserve protection, but who do you think that you are protecting when you protect languages? Exactly, when you protect a language you are protecting people who speak it. If you want to get rid of your native language, fine, but do not force other people to do so.

I don't want to force anyone to do anything. I just express my opinion. And certain things are just facts like learning English will give you superior economical power.

I don't think there is any reason to protect language from criticism. Polish declination an conjugation as well as use of a lot "sz", "cz", "dz" etc. makes it hard to learn language. It's a fair criticism. If you want to do machine learning in Polish you will have a lot of issues. No one should get offended by that.

Also, there is also a lot of forms that makes Polish difficult to use for native speakers like "wyszedłem" (I, as a man, went out) the correct spelling vs. "wyszłem" an incorrect spelling that seems to be easier to use for many people.

Therefore, a lot of language protection is focused on excluding people. You don't know how to use language correctly so you are inferior. It's very common sentiment among Polish people. We even have a special body to say what should be considered correct - Polish Language Council. French have Académie française.

Even in English, especially in academia, people like to have this snobbish attitude that you should not mix British and American forms in a single text. This is certainly not intended to protect anyone, but to exclude people less proficient in English.

I think the most stark example of excluding people based on language I saw myself was in Belgium in Dutch speaking Flanders. If I'm not mistaken they have a law that explicitly forbids teaching bachelor degrees in anything but Dutch. This excludes a lot of French speaking Walloons from participating in Flemish universities. Some of them end up in Netherlands where attitude to language is less protectionists and a lot of bachelor programs are in English.

On the other hand, in Wallonia you may come across people who will pretend to not speak English. They will understand you, but they will be hostile. It's their part of country, their rules, but I'm not afraid to call it rude.

Again, people deserve protection. For example, you should be able to conduct your business with your government in the language that you speak. But there is no need to be protective of languages themselves and often this protectionism is a weapon on its own.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#116
post #114
post #97

For those who want to know more about this: The article touches on just the tip of the iceberg. You might think that all you need to do is add an extra field for phonetic readings, and then simply sort on that field, but there are a lot of things that can go wrong. A naive sort (i.e. based simply on character code) will hit the following snags: 1) Hiragana vs Katakana The article focuses more on Kanji vs Kana, but Ja…

I'd assume you'd want ゅ andっ to be added to the kana they're attached to, and then sorted. I'd want my き and きゅ together and び and っび together.

That does make sense in a way, but I don't think that would feel natural to any native Japanese speaker. I'm not native and even I would find that ordering very odd.

At the very least, it would make the sorting algorithm a lot more complex if you had to look ahead at later characters in order to sort the current prefix.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#117
post #20
post #3

https://en.m.wikipedia.org/wiki/Four-Corner_Method You can sort Chinese characters (including Kanji but i'm not sure they use the Four Corners Method) by the Four Corners method. Why would you need to sort kanji phonetically in the first place? Do Japanese users actually expect names to be sorted phonetically? English speakers don't expect names to be sorted by IPA so consistency of the sorting scheme should be all t…

> Do Japanese users actually expect names to be sorted phonetically? Yes - or that's how a human would sort them, at any rate.

Think about numbers, it is not sorted phonetically, Even the alphabet is not. All those languages that use alphabet sort in exactly the same order regardless the pronunciation.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#118
post #93
post #20

Earlier quoted context omitted.

> Do Japanese users actually expect names to be sorted phonetically? Yes - or that's how a human would sort them, at any rate.

I don't think you're in a position to speak for all of humanity.

I meant a human as opposed to an algorithm.

Point being: kanji words aren't always sorted phonetically, for the reasons described in the article, so a user may not be surprised if they aren't. But when a human is sorting kanji words they do so phonetically by reading.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#119

Am I missing something subtle in this Kanji example, or all 4 names actually written the same?: "There are four Japanese women whose names you have to sort: Junko, Atsuko, Kiyoko, and Akiko. This does not seem difficult, until they each show you how they write their names in kanji: 淳子 (Junko) 淳子 (Atsuko) 淳子 (Kiyoko) 淳子 (Akiko)" I'm not familiar with Japanese at all (and have never had to deal with localization beyond…

On Japanese forms that require you to use your name, you provide them with how your name looks in kanji, as well as how they are read (in their syllabic alphabet). When introducing yourself in speaking, you may also mention how your name is written. It's not a problem in the sense that when you are in a position to ask someone for their name, you are also in a position to ask them for both the orthographic and phonet…

Thanks, this was very informative.

Japanese, kanji in particular, seems to me to be unnecessarily complex (from an American that's spent most of hos life in the boonies).

You mention they have a separate phonetic alphabet to describe kanji? I know I'm barking up the wrong tree, but why not abandon kanji and just use the phonetic alphabet? Seems to get to where everyone wants to be and without the ambiguity that kanji involves.

I imagine the answer to these probably comes down to some combination of tradition or cultural pride. I don't mean this with any malice, it's just my curiosity and likely cultural ignorance.

Re: Sorting in Japanese – An Unsolved Problem (2011)

#120

Earlier quoted context omitted.

On Japanese forms that require you to use your name, you provide them with how your name looks in kanji, as well as how they are read (in their syllabic alphabet). When introducing yourself in speaking, you may also mention how your name is written. It's not a problem in the sense that when you are in a position to ask someone for their name, you are also in a position to ask them for both the orthographic and phonet…

Thanks, this was very informative. Japanese, kanji in particular, seems to me to be unnecessarily complex (from an American that's spent most of hos life in the boonies). You mention they have a separate phonetic alphabet to describe kanji? I know I'm barking up the wrong tree, but why not abandon kanji and just use the phonetic alphabet? Seems to get to where everyone wants to be and without the ambiguity that kanji…

> You mention they have a separate phonetic alphabet to describe kanji? I know I'm barking up the wrong tree, but why not abandon kanji and just use the phonetic alphabet?

It is tradition + cultural pride, and cost of switching. It's worth noting that in the modern era, governments of almost all countries that traditionally used Chinese characters for writing undertook massive reforms to limit their use.

* In Vietnam, French colonization forcibly replaced the traditional education system, and along with that introduced the Latin script for writing.

* In Korea, Hangul was invented by one of its kings, with the letterforms created from an abstract representation of the oral organs that are exercised in creating each particular sound, nevertheless, the script it only really took off centuries later alongside a growing Korean nationalism.

* In Japan, a phonetic alphabet developed that can be written in two styles (think capital/small letters); one from components of Chinese characters, the other from cursive writing of those characters. There were certainly proposals to eliminate Chinese characters, but what happened was that the number of Chinese characters in publications that could be used in publications without being accompanied by a phonetic spelling was limited to a couple thousand, and this persists to this day. Note that Japanese names are allowed to draw from a much larger repertoire of characters.

* In communist China, under communist influence, the government embarked on a project of simplifications of Chinese characters, with possibly an objective of reforming it to a phonetic alphabet. The first round of simplifications were based on common shorthand and cursive styles. The second round incorporated more drastic simplifications with less precedent, and it failed miserably, and the Chinese abandoned the project.

* Hong Kong and Taiwan did not seem to have embarked on any simplification project. There probably exists some pride at continuing to use traditional characters.

In all these countries, resistance to reform came from the literati, for which knowledge of Chinese characters were a shibboleth for elitism. Two arguments are frequently trotted out:

* One is that Chinese characters helped to distinguish homonyms. But investigating the situation in Vietnam and Korea, this is likely to be quite unimportant in clear communication. I suppose that this is because context will make clear in most instances, and use of synonyms should make up for the others.

* Another is that a change of script would cut one off from their existing literature. But already, the respective languages have changed so much that modern readers cannot read ancient texts without a lot of guidance and learning new meaning.

Interestingly, in modern Japan, there exists an problem with kanji familiarity, and I think this phenomenon is growing in China too. More and more people are forgetting how to write Chinese characters, even though they can recognize them perfectly. How one enters kanji into computers is by entering the phonetic pronunciation, and using computer software to identify the correct characters to convert them to. So you can see that one no longer creates the character from memory very much anymore, and when you don't practice it you lose it.

If you look at historical context, the reforms were always associated with westernization and nationalism. I think that if Japan were to seriously consider reform to a completely phonetic script, it would be by associating with the phonetic script a notion of national pride. In this respect the Japanese phonetic scripts hold less promise than Hangul, for while the former is a derivative of chinese letterforms, the latter was created by a national king, no less, and based rationally on depicting fundamental linguistic phenomena.

P.S. y'all Americans should adopt metric.

Post reply on HN