Live data from Hacker News

I Can’t Write My Name in Unicode

modelviewculture.com

311–320 of 377 posts

Re: I Can’t Write My Name in Unicode

#311

Earlier quoted context omitted.

Neither of your examples are solveable by the Unicode consortium, who are not the god-emperors of fonts (nor would you want them to be).

I respectfully disagree. If Japanese ideograms and Chinese ideograms actually used different code points (i.e. no "Han unification"), then the problem wouldn't exist - the phone could trivially use a Japanese font for Japanese text, and a Chinese font for Chinese text.

No. Using different code points for the same character used in different languages creates big problems. It would be like having different code points for 'A' depending on whether it was used in English, Spanish, German, etc. If you somehow ended up writing "color" with both 'o' characters from the Spanish ABCs and the others from the English ABCs, you'd have a real mess when it came to sorting, searching, name matching (what language is "Hans"?) etc. It is far more convenient to allow the character sequence "color" or "Hans" to be language independent, even if the font choices, pronunciation, sort order, etc., are language dependent.

Chinese, Japanese, and Korean writers face similar issues. The characters they use to write the name of China or Japan, the ten digits, the characters for year, month, and day in dates, and so many thousands of others in Chinese characters are what they all consider to be the same characters. That is not all characters, but it includes so many that insisting on different code points by language would make a real mess. Hong Kong has many characters that are unique to HK Cantonese. So, should Cantonese have a full set of all Chinese characters that are the "Cantonese characters"? How about Shanghainese, then? Or Hakka? Teochew (Chaozhou) or a dozen Chinese languages? Full, independent sets of all Chinese characters for each? Suppose you accidentally used an input method in HK and wrote the name of some Beijing gov't ministry using characters that looked identical to their Mandarin counterparts but were entirely different codepoints? Now what? You can't find your search term? You mess up the database and have two identical-looking keys?

No, Han unification is not conceptually different from unifying ABCs used by English and Spanish speakers, Cyrillic used by Russians or Serbs, etc., except that there are many more characters, so the boundary between what should be unified and what shouldn't contains more items in the gray zone to cause debate. Having no Han unification at all wouldn't solve all problems, it would create all sorts of absurdity.

Re: I Can’t Write My Name in Unicode

#312

Earlier quoted context omitted.

One by one: I'm confused by your use of 'our' and 'we'. It seems you're trying to write from the general point of view of a German, answering .. a German? Are umlauts letters? Yes. [1] [2] Maybe not the best source, but please provide a better one if you disagree so that I can actually understand where you're coming from. I understand - I hope? - composition. And I tend to agree that it shouldn't matter much if the i…

I am not entirely sure if Germans count umlauts as distinct characters or modified versions of the base character. And maybe it is not so important; they still do deserve their own code points. Note BTW that in e.g. Swedish and German alphabets, there are some overlapping non-ASCII characters (ä, ö) and some that are distinct to each language (å, ü). It is important that the Swedish ä and German ä are rendered to the…

Umlauts are not distinct characters, but modifications of existing ones to indicate a sound shift.

http://en.wikipedia.org/wiki/Diaeresis_%28diacritic%29

German has valid transcriptions to their base alphabet for those, e.g "Schreoder" is a valid way to write "Schröder".

ß, however, is a separate character that is not listed in the german alphabet, especially because some subgroups don't use it. (e.g. swiss german doesn't have it)

Re: I Can’t Write My Name in Unicode

#313
post #117

Earlier quoted context omitted.

> The Bengali fonts that you have installed happen to render it as a jophola It's not only the Bengali font - the text rendering framework of my operating system also needs to have a bunch of complex rules to figure out that a jophola needs to be rendered. It also needs to know that the visual ordering of i-kar is before the preceding consonant cluster (দ in আদিত্য). > the characters that are required to type a jopho…

Is Bengali your first language? While one can make the case that ত্য is simply "'to' - 'o' + 'ya' = 'to'"[0][1], it's rather confusing mental acrobatics, and it doesn't reflect either how the writing system is taught, or how native speakers use it and think of it on a day-to-day basis. If anything, your comment makes a stronger argument for consolidating ই and ি (they are literally the same letter and phoneme, but wr…

Is Bengali your first language?

A better question is, Are there any native Bengali speakers creating character set standards in Bangladesh or India? If not, why not? If so, did they omit your character?

I ask, because although you prefer to follow the orthodox pattern of blaming white racism for your grievance du jour, the policy of the Unicode Technical Committee for years has been to use the national standards created by the national standards bodies where these scripts are most used as their most important input.

Twenty years ago, I spent a lot of time in these UTC meetings, and when the question arose as to whether to incorporate X-Script into the standard yet, the answer was never whether these cultural imperialists valued, say, Western science fiction fans over irrelevant foreigners, but it was always, "What is the status of X-Script standardization in X-land?" Someone would then report on it. If there was a solid, national standard in place, well-used by local native speakers in local IT applications, it would be fast-tracked into Unicode with little to no modification after verification with the national authorities that they weren't on the verge of changing it. If, however, there was no official, local standard, or several conflicting standards, or a local standard that local IT people had to patch and work around, or whatever, X-Script would be put on a back burner until the local experts figured out their own needs and committed to them.

The complaint in this silly article about tiny Klingon being included before a complete Bengali is precisely because getting Bengali right was more complex and far more important. Apparently, the Bengali experts have not yet established a national standard that is clear, widely implemented, agreed upon by Bengali speakers and that includes the character the author wants in the form he/she wants it, for which he/she inevitably blames "mostly white men."

(Edited to say "he/she", since I don't know which.)

Re: I Can’t Write My Name in Unicode

#314
post #48

Not sure if the l33tspeak analogy is fully justified. In case of the "missing" letter (called khanda-ta in Bengali) for the Bengali equivalent of "suddenly", historically, it has been a derivative of the ta-halant form (ত + ্ + ‍ ). As the language evolved, khanda-ta became a grapheme of its own, and Unicode 4.1 did encode it as a distinct grapheme. A nicely written review of the discussions around the addition can b…

FWIW there is an interesting project called Swarachakra[1] which tries to create a layered keyboard layout for Indic languages (for mobile) that's intuitive to use. I've used it for Marathi and it's been pretty great.

They also support Bengali, and I bet they would be open to suggestions.

[1]: http://en.wikipedia.org/wiki/Swarachakra

Re: I Can’t Write My Name in Unicode

#315

“Whatever path we take, it’s imperative that the writing system of the 21st century be driven by the needs of the people using it. In the end, a non-native speaker – even one who is fluent in the language – cannot truly speak on behalf the monolingual, native speaker.” Not sure how the author can simultaneously say this, while criticizing the CJK unification, which makes total sense, and has never been a point of con…

Actually yes, the CJK unification is a problem for many people, including me when I want to read Japanese on a phone bought in Europe. Example 1. Typically, any time you want to mix the 2 languages you're getting in trouble. Let's say you write a textbook for Chinese people to learn Japanese as a second language. Or a research article in Japanese citing old Chinese literature. In your text, you'll have to mark specif…

> In your text, you'll have to mark specifically which part are in Japanese and which part are in Chinese and use different font for them. If you don't, the characters look wrong.

Most non-trivial text-rendering systems need to know the language of the text in order to properly render it. This is not only an issue for Japanese/Chinese but also e.g. English/German with different hyphenation rules[0], different rules for ligatures[1] and different rules/preferences for spacing after punctuation[2]. None of these can sensibly be solved by Unicode.

[0] ‘backen’ (to bake) becomes bak-ken when hyphenated at the end of a line, ‘airlock’ can’t be hyphenated at all between c and k.

[1] ‘Schiff’ (ship) should be rendered with a ligature ff, ‘Dampfflut’ (~ steam flow?) shouldn’t. ‘afferent’ probably should have it.

[2] Single space after ‘.’ and ‘,’ in German, whereas French tends to prefer more space after ‘.’.

Re: I Can’t Write My Name in Unicode

#316
post #313

Earlier quoted context omitted.

Is Bengali your first language? While one can make the case that ত্য is simply "'to' - 'o' + 'ya' = 'to'"[0][1], it's rather confusing mental acrobatics, and it doesn't reflect either how the writing system is taught, or how native speakers use it and think of it on a day-to-day basis. If anything, your comment makes a stronger argument for consolidating ই and ি (they are literally the same letter and phoneme, but wr…

Is Bengali your first language? A better question is, Are there any native Bengali speakers creating character set standards in Bangladesh or India? If not, why not? If so, did they omit your character? I ask, because although you prefer to follow the orthodox pattern of blaming white racism for your grievance du jour, the policy of the Unicode Technical Committee for years has been to use the national standards crea…

I mostly agree with your point, but note that the author is male (well, the name is a commonly male one).

It's a bit telling that folks in the software industry[1] seem to assume that techies are male (a priori), but those who write articles of this kind are female.

Not blaming you for it, but it's something you should try to be conscious about and fix.

[1] I've been guilty of this myself, though usually in cases where I use terms like "guys" where I shouldn't be.

Re: I Can’t Write My Name in Unicode

#317
post #77
post #72

Earlier quoted context omitted.

I don't understand Bengali at all. I'm trying to understand your second sentence though. When you say "no one writes that way", do you mean nobody hits the keys for letter, followed by vowel-silencing diacritic, followed by another vowel? Or do you mean the glyph that results from that combination of keystrokes doesn't match how a Bengali speaker would write it on paper? If it's the latter, isn't that an issue for th…

It has to do with how the text is rendered. For example, if you see the Bengali text on page 3 of this PDF: http://www.unicode.org/L2/L2004/04252-khanda-ta-review.pdf it is unreadable and incorrect Bengali. ;-)

Unicode is about _code points_, not what is actually rendered. It recommends what should be shown, but it's up to the font on how to render it.

Re: I Can’t Write My Name in Unicode

#318
post #55

Earlier quoted context omitted.

I don't understand Bengali at all, but the character ৎ does have it's own Unicode codepoint (U+09CE, BENGALI LETTER KHANDA TA). It was introduced in Unicode 4.1 in 2005.

Yes, this is what the article says: > Until 2005, Unicode did not have one of the characters in the Bengali word for “suddenly”. The codepoint does exist now, but it took ten years before it was included.

... and emoji was added to the Unicode standard in 2010. The title of your article is fallacious :/

(though I agree with most of it)

Re: I Can’t Write My Name in Unicode

#319

Earlier quoted context omitted.

I am not entirely sure if Germans count umlauts as distinct characters or modified versions of the base character. And maybe it is not so important; they still do deserve their own code points. Note BTW that in e.g. Swedish and German alphabets, there are some overlapping non-ASCII characters (ä, ö) and some that are distinct to each language (å, ü). It is important that the Swedish ä and German ä are rendered to the…

Umlauts are not distinct characters, but modifications of existing ones to indicate a sound shift. http://en.wikipedia.org/wiki/Diaeresis_%28diacritic%29 German has valid transcriptions to their base alphabet for those, e.g "Schreoder" is a valid way to write "Schröder". ß, however, is a separate character that is not listed in the german alphabet, especially because some subgroups don't use it. (e.g. swiss german do…

Yes, that transcription approach is familiar; here the result of German-Swedish-Finnish equivalency of "ä" is sometimes not so good.

For instance, in skiing competitions, the start lists are for some reason made with transcriptions to ASCII. It's quite okay that Schröder becomes Schroeder, but it is less desirable that Söderström becomes Soederstroem and quite infuriating that Hämäläinen becomes Haemaelaeinen. We'd like it to be Hamalainen, just drop the dots.

Re: I Can’t Write My Name in Unicode

#320
post #66

Earlier quoted context omitted.

> It's just internally represented as multiple codepoints And in fact it is not, and even in the article it is U+09CE. One codepoint. If his input method irks him, he's as free to tweak it as I am to switch to Dvorak. Also folks, there's no "CJK unification" project. It's Han unification. Han characters are Han characters, just like Latin characters are Latin characters. Just because German has ß and Danish has Ø doe…

I hate to say it, but I think the author's objections seem to stem from his lack of understanding of character encoding issues. I don't know Bengali at all and so I will try to refrain from commenting on it, but I do speak and read Japanese fluently and Han Unification is a very, very good thing . Can you imagine the absolute hell you would have to go through trying to determine if place names were the same if they u…

I feel that way too. The distinction between codepoint, glyph, grapheme, character, (...) is not an easy one, and that's what he seems to be stumbling over. Unicode worries itself about only some of these issues, many of the other issues are about rendering (the job of the font) or input.
Post reply on HN