Live data from Hacker News

I Can’t Write My Name in Unicode

modelviewculture.com

161–170 of 377 posts

Re: I Can’t Write My Name in Unicode

#161
post #48

Not sure if the l33tspeak analogy is fully justified. In case of the "missing" letter (called khanda-ta in Bengali) for the Bengali equivalent of "suddenly", historically, it has been a derivative of the ta-halant form (ত + ্ + ‍ ). As the language evolved, khanda-ta became a grapheme of its own, and Unicode 4.1 did encode it as a distinct grapheme. A nicely written review of the discussions around the addition can b…

> I could write the author's name fine: আদিত্য Author here. Well, yes and no. The jophola at the end is not actually given its own codepoint[0]. The best analogy I can give is to a ligature in English[1]. The Bengali fonts that you have installed happen to render it as a jophola, the way some fonts happen to render "ff" as "ff" but that's not the same thing as saying that it actually is a jophola (according to the Uni…

> whereas the characters that are required to type a jophola have no obvious semantic, phonetic, or orthographic connection to the jophola.

Then it is an input method issue not an encoding issue.

Re: I Can’t Write My Name in Unicode

#162
It is much easier to criticize than to fix it.

While it is good to bring awareness to this, we are still growing in this area. In fact we should applaud the efforts so far that we even have a standard that somewhat works for most of the digital world. Does it need to evolve further, yes.

I am sure the engineers and multi-lingual people that stepped up to do Unicode and organize it aren't trying to exclude anyone. Largely it comes down to who has been investing in the progress and time. It may even be easier to fund and move this along further in this time, it was hard to fund anything like this before the internet really hit and largely relied on corporations to fund software evolution and standards.

In no way should the engineers or group getting us this far (UC) be chided or lambasted for progressing us to this step, this is truly a case of no good deed goes unpunished.

Re: I Can’t Write My Name in Unicode

#163

> He proudly announces that there are ‘no fewer than 147 Indian dialects’ – a pathetically inaccurate count. Wow. How can a country function like this? Is everyone proficient in their native language plus a 'common' one, or are all interactions supposed to be translated inside the same country? Regardless of historical and cultural value, if that's the case, it seems... inefficient. I do realize that there are more c…

> I do realize that there are more countries like this,

"More"? I wouldn't be surprised if it was most.

Re: I Can’t Write My Name in Unicode

#164
> No native English speaker would ever think to try “Greco Unification” and consolidate the English, Russian, German, Swedish, Greek, and other European languages’ alphabets into a single alphabet.

This actually is a pretty good idea. Cyrillic, Latin, and other "Greco" scripts share quite a lot of characters. There's no need for both А (http://www.fileformat.info/info/unicode/char/0410/index.htm) and A (http://www.fileformat.info/info/unicode/char/0041/index.htm) beyond ASCII and other legacy compatibility.

Re: I Can’t Write My Name in Unicode

#166

I came in expecting to read an article bemoaning some niche language and playing the diversity card. I was not disappointed, but as I kept reading, the author made some very good points. I don't really care that the organization is run by white men who speak English, because frankly the entire computing industry and telecommunications industry is based on that. I'm not going to argue about the original sin there, bec…

> Personally, I think that it's nice to have an ever-decreasing number of languages to support, preferably English. The first part of that is because it's annoying enough to unambiguously parse one human language much less dozens, and the second part is pure convenience and small-mindedness on my part.

I'll go further and claim that it's all about convenience and small-mindedness.

It's so typical of a certain group of programmers to complain like this. Ugh, why do we have different languages and ways of doing things? Let's just converge on the one I grew up with, or use[1]. Hopefully you won't be forced to work on this kind of thing, or any kind of work that you don't care about/think is worth it. But until you do, maybe you could keep your self-centred navel gazing to yourself.

[1] In the case of English, most programmers who express this opinion write perfectly good English, though they might not be native speakers. They wouldn't be inconvenienced at all if suddenly everyone forgot about their own language and started speaking English, especially since they don't seem to care about languages beyond having some way of communicating with other people.

Re: I Can’t Write My Name in Unicode

#167
post #155

Earlier quoted context omitted.

Is Bengali your first language? While one can make the case that ত্য is simply "'to' - 'o' + 'ya' = 'to'"[0][1], it's rather confusing mental acrobatics, and it doesn't reflect either how the writing system is taught, or how native speakers use it and think of it on a day-to-day basis. If anything, your comment makes a stronger argument for consolidating ই and ি (they are literally the same letter and phoneme, but wr…

> Is Bengali your first language? Yes. > [...] it's rather confusing mental acrobatics, and it doesn't reflect either how the writing system is taught, or how native speakers use it and think of it on a day-to-day basis. Mental acrobatics are part-and-parcel of the language, either in digital or non-digital form. If I were to spell out your name aloud, I would end with "ত-এ য-ফলা", which doesn't really say anything a…

> My point in the original comment (and to some extent in the preceding one) was to emphasize that a lot of these issues are at the input method level - we should not have to think about encoding as long as it accurately and unambiguously represent whatever we want it to represent.

I might be sympathetic to this, except that keyboard layouts and input (esp. on mobile devices) is an even bigger mess and even more fragmented than character encoding. Furthermore, while keys are not a 1:1 mapping with Unicode codepoints, they are very strongly influenced by the defined codepoints.

It'd be nice to separate those two components cleanly, but since language is defined by how it's used, this abstraction is always going to be very porous.

> Just out of curiosity - I would be interested to know more about your learning experience that you feel is not well aligned with the representation of jophola as it is currently.

I have literally never once heard the jophola referred to as a viram and a য, except in contexts such as this one. Especially since the jophola creates a vowel sound (instead of the consonant য), and especially since the jophola isn't even pronounced like "y". (I understand why the jophola is pronounced the way it does, but arguing on the basis of phonetic Sanskrit is a poor representation of Bengali today - by that point, we might as well be arguing that the thorn[0] is equivalent to "th" today, or that æ should be a separate letter and not a dipthong[1])

I'm not going to claim that it's completely without pre-digital precedent, but it certainly is not universal, and it's inconsistent in one of the above ways no matter how one slices it, especially when looking at some incredibly obscure and/or antiquated modifiers that are given their own characters, despite being undeniably joined of other characters that are already Unicode codepoints[2].

(Out of curiosity, where in Bengal are you from?)

[0] https://en.wikipedia.org/wiki/Thorn_%28letter%29

[1] As it was in Old English.

[2] Such as the ash (æ)!

Re: I Can’t Write My Name in Unicode

#168

Earlier quoted context omitted.

Imagine if the letter Q had been left out of Unicode's Latin alphabet. The argument against it is that it can be written with a capital O combined with a comma. (That's going to play hell with naive sorting algorithms, of course, but oh well.) Oh, and also imagine your name is Quentin.

My first choice as theoretical Quentin wouldn't be "how can I frame this accidental, perhaps even flagrantly disrespectful omission as antiprogressive and dissect the credentials, experience, and ethnicity of the people who made the mistake via culture essay," it would probably be "where do I issue a pull request to fix this mistake or in what way can I help?" Maybe that's just me. I look forward to the future where…

> …including a weekly "worst of HN" comment dissection.

That sounds interesting, but I can't find any reference to it on their site or search engines. Do you have a link?

Re: I Can’t Write My Name in Unicode

#169
post #147
post #143

Earlier quoted context omitted.

This is the core of the Han Unification debate. "G" and "g" are the same letter. They started off as stylistic forms of a unicameral alphabet. Over time they took on separate meanings, and now we have a bicameral alphabet, where the two forms have different code points. Of course, over the last 2000 years, we've developed rules for how to use them. "I was reading a nice book on Polish polish on the way from Reading t…

Absolutely. But the two "g" haven't diverged, yet. We don't give out code points to speculative future developments. If and when they diverge one will get its very own code point.

The premise of Someone's hypothetical was to explain to theon144 why it might be both hard and important to guess. The hypothetical assumed that the difference already existed. It echos the larger context of Han unification that com2kid started, but with an example that's a lot easier for native English speakers to understand.

So while I agree that they haven't yet diverged, that's outside the context of the hypothetical, where they have diverged.

Re: I Can’t Write My Name in Unicode

#170

“Whatever path we take, it’s imperative that the writing system of the 21st century be driven by the needs of the people using it. In the end, a non-native speaker – even one who is fluent in the language – cannot truly speak on behalf the monolingual, native speaker.” Not sure how the author can simultaneously say this, while criticizing the CJK unification, which makes total sense, and has never been a point of con…

Not to say there aren’t problems with CJK unification. It’s difficult or impossible to represent old family names or newly coined characters without some means of composing characters from radicals—even if you do want “precomposed” characters a majority of the time, as is the case with, say, é (00E9) versus é (0065 0301). As far as these characters being “the same”, I think it’s better to say they’re analogues . A i…

> I would actually support the author’s “Greco Unification” strawman if it could be done in a principled way

A 5 year old tech note by a Unicode Consortium member at http://www.unicode.org/notes/tn26 gives 7 reasons "why the Latin, Greek, and Cyrillic scripts have been separately encoded, rather than being encoded as a single script", and 12 reasons why the Han script was unified.

Edit: These 7 reasons make a good case why such "Greco Unification" wouldn't work.

Post reply on HN