Earlier quoted context omitted.
no, it isn't a more apt comparison. Antiqua and Fraktur look much more different from each other than most Chinese and Japanese characters are from eachother. Here are two screenshots in Japanese and Chinese, taken from the http://www.yomiuri.co.jp/ for one, and http://www.cntv.cn/ for the other. Neither media outlet can be accused of being cultural sellouts misrepresenting the written tradition of their countries. J…
Your screenshots are showing the exact same glyphs because you don't have the "宋体" font installed that's specified by the Chinese web page (it's a script font so you'd notice instantly it's different), so your browser picked a Japanese font to render it instead. The only reason they look the same in your screenshots is due to Han Unification. Your reasoning is "they're unified in Unicode which is proof they should be…
I Can’t Write My Name in Unicode
341–350 of 377 posts
Re: I Can’t Write My Name in Unicode
#342Earlier quoted context omitted.
I respectfully disagree. If Japanese ideograms and Chinese ideograms actually used different code points (i.e. no "Han unification"), then the problem wouldn't exist - the phone could trivially use a Japanese font for Japanese text, and a Chinese font for Chinese text.
But.. you already posted a solution to your problem - use chinese font for chinese and japanese font for japanese. There is no problem. I mean, you would also presumably use a western font for latin and another separate font for cyrillic, since japanese fonts universally have absolutely dreadful kerning on latin (and often omit cyrillic entirely).
Re: I Can’t Write My Name in Unicode
#343Earlier quoted context omitted.
Actually, that's not correct, and it's the exact same mistake I made when using that API. codePointAt returns the codepoint at index i, where i is measured in 16-bit chars, which means you could index into the middle of a surrogate pair. The correct version is: for (int i = 0; i Java 8 seems to have acquired a codePoints() method on the CharSequence interface which seems to do the same thing. But this just adds to th…
I think you missed the part where `i` is not incremented in the for statement, but inside the loop using `Character.charCount`, which returns the number of `char` necessary to represent the code point. If there's something wrong with this, my unit tests have never brought it up, and I am always sure to test with multi-`char` codepoints.
Re: I Can’t Write My Name in Unicode
#344Earlier quoted context omitted.
English is the best candidate because it has the second largest user base (1.2 Billion vs 1.3 Billion for Mandarin), http://en.wikipedia.org/wiki/List_of_languages_by_total_numb... and is twice as spoken as the third most popular language Spanish. (0.55 Billion) If I got to pick the universal language, it would be Lojban (a few hundred speakers), but that is not a realistic goal, teaching the other 6 Billion people a…
Even if 1.2 billion seems a lot, that's still a small fraction of a world's population. So every choice of a universal language would force majority of a world to learn new one. So that's why I think winning popularity contest is a poor argument and we shouldn't look at that and focus on things like simplicity (which I don't find in English), speed of learning, consistency, expressiveness etc. I'd be happy to use Loj…
Re: I Can’t Write My Name in Unicode
#345Earlier quoted context omitted.
It applies to Ä and Æ... which is what the parent said. It doesn't apply to 気 and 气 and 氣. Those are all the same thing.
気 气 and 氣 are actually not merged by han unification, and could be described as similar to similar to Ä and Æ: shared etymology, same meaning, some languages decided to use a simpler character because the old one was too complicated to write. The variations on characters which have been merged are usually even closer than that. More like the single or double storey "a", or the single or double loop "g".
I think what you're missing is that Han unification never particularly concerned itself with whether characters are "the same" or not. Somebody basically just decided that some differences were important and others weren't. I mean, the difference between 語 and 语 is analogous to printing and cursive, but for whatever reason it made the cut. Meanwhile the SC and TC versions of 骨 are approximately mirror images, but to show you I'd need two different fonts.
Anyway the merged differences are not just analogous to different fonts.
Re: I Can’t Write My Name in Unicode
#346Earlier quoted context omitted.
I respectfully disagree. If Japanese ideograms and Chinese ideograms actually used different code points (i.e. no "Han unification"), then the problem wouldn't exist - the phone could trivially use a Japanese font for Japanese text, and a Chinese font for Chinese text.
No. Using different code points for the same character used in different languages creates big problems. It would be like having different code points for 'A' depending on whether it was used in English, Spanish, German, etc. If you somehow ended up writing "color" with both 'o' characters from the Spanish ABCs and the others from the English ABCs, you'd have a real mess when it came to sorting, searching, name match…
It's difficult to come up with a logical explanation for why European languages that use their own alphabet get their own codepoints, but ideographic languages need to be "unified", even though the actual letters as used in those languages look different.
The "Han unification" was fundamentally a bad idea, and persists for historical reasons. Back when (some) people thought a fixed-width 16-bit character representation would be "enough", it made sense to try to reduce the number of "redundant" code points. Now that Unicode has expanded to a much-larger code space, I would think they'd choose differently.
Unfortunately, that kind of sweeping change is unlikely any time soon.
Re: I Can’t Write My Name in Unicode
#347Re: I Can’t Write My Name in Unicode
#348The argument that author makes is "every letter in English alphabet is represented, why not every letter/grapheme in Bengali/Tamil/Telugu/Name your language" argument is specious at best
Except that the whole purpose of Unicode is to create a character encoding that "enables people around the world to use computers in any language" - taken from the Unicode Consortium website. Bengali is also not an obscure language. It is the 10th most spoken language in the world and the national language of Bangladesh.
Re: I Can’t Write My Name in Unicode
#349Earlier quoted context omitted.
This is a terrible excuse: the Unicode Consortium should always seek out at least one (if not a group of) native speakers of a language before defining code points for that language. These speakers really should be both native speakers of and experts in the language. There are countless ways to reach out to Bengali speakers, only one of which is the Indian government - whatever politics governments may play, a techno…
That's ridiculous. How can you expect UC to reach out to the 1000s of languages out there? Plus, shouldn't the party who expects to benefit put in the effort? Did you know that the Indian Government has a department for just such a thing, http://tdil.mit.gov.in/ ? What has it been doing all these years? Added later: there's also CDAC http://cdac.in/index.aspx?id=mlingual . And these are just 2 that I found quickly.
Re: I Can’t Write My Name in Unicode
#350Earlier quoted context omitted.
My first choice as theoretical Quentin wouldn't be "how can I frame this accidental, perhaps even flagrantly disrespectful omission as antiprogressive and dissect the credentials, experience, and ethnicity of the people who made the mistake via culture essay," it would probably be "where do I issue a pull request to fix this mistake or in what way can I help?" Maybe that's just me. I look forward to the future where…
"My first choice as theoretical Quentin wouldn't be "how can I frame this accidental, perhaps even flagrantly disrespectful omission as antiprogressive and dissect the credentials, experience, and ethnicity of the people who made the mistake via culture essay," it would probably be "where do I issue a pull request to fix this mistake or in what way can I help?" " I could tell you, but I'll need $18,000 first.