Live data from Hacker News

I Can’t Write My Name in Unicode

modelviewculture.com

51–60 of 377 posts

Re: I Can’t Write My Name in Unicode

#51
post #4

I wonder if the author has submitted a proposal to get the missing glyph for their name added. You don't need to be a member of the consortium to propose adding a missing glyph/updating the standard. The point of the committee as I understand it isn't to be an expert in all forms of writing, but to take the recommendations from scholars/experts and get a working implementation, though more diverse representation of l…

It sounds like the glyph is in unicode already, but expressed using combining characters?

His writing there is pretty confusing. He started by complaining about a glyph that was missing until 2005, but was either fixed in 2005 or approximated by combining some characters 'ত + ্ + ‍ = ‍ৎ'. He doesn't really make it very clear whether ৎ is a substitute for a glyph, or whether it's the correct glyph and a case of an input system not making that easy to enter, but it seems like the glyph was added in 2005 and he's complaining about the input method. Assuming it is a case of a clunky input system, then pointing the finger at the Unicode consortium seems pretty weak, since so far as I understand it, various OS vendors/app platforms handle that implementation.

Re: I Can’t Write My Name in Unicode

#52

Earlier quoted context omitted.

That's ridiculous. How can you expect UC to reach out to the 1000s of languages out there? Plus, shouldn't the party who expects to benefit put in the effort? Did you know that the Indian Government has a department for just such a thing, http://tdil.mit.gov.in/ ? What has it been doing all these years? Added later: there's also CDAC http://cdac.in/index.aspx?id=mlingual . And these are just 2 that I found quickly.

"The Unicode Consortium is a non-profit corporation devoted to developing, maintaining, and promoting software internationalization standards and data, particularly the Unicode Standard, which specifies the representation of text in all modern software products and standards." The fact that it is their goal to set the standard for the textual representation of human speech means that they take on that responsibility.

This is like complaining that Wikipedia has only 35K Bengali articles, while English has over 4.7M!

Why is Wikipedia not doing more to translate the 4.7M articles into Bengali? It does claim to be the world's encyclopedia!

Re: I Can’t Write My Name in Unicode

#53

Earlier quoted context omitted.

I found it extremely annoying that he doesn't even say specifically what the problem with his name is . So there's a letter that's unavailable? Which letter?

> Even today, I am forced to do this when writing my own name. My name is not only a common Indian name, but one of the top 1,000 names in the United States as well. But the final letter has still not been given its own Unicode character, so I have to use a substitute. Not as descriptive as it could be, but this article isn't about him.

Yet the title is solely about him.

Re: I Can’t Write My Name in Unicode

#54

“Whatever path we take, it’s imperative that the writing system of the 21st century be driven by the needs of the people using it. In the end, a non-native speaker – even one who is fluent in the language – cannot truly speak on behalf the monolingual, native speaker.” Not sure how the author can simultaneously say this, while criticizing the CJK unification, which makes total sense, and has never been a point of con…

> [...] CJK unification[...] has never been a point of contention in the communities concerned with it.

I am not very familar with the CJK unification project so take my points with a grain of salt.

> more than not opposing CJK unification, I benefit from it greatly.

I think think that is a different point of view. Isn't it? You are seeing your benefit whereas the author is seeing his. Here's an alternative solution: What if the search engine understood what you were searching for and returned results in all the languages? Unification can result in a lot of information loss the same way a photo can be compressed but it comes at the cost of loss in quality.

> so there is(to my eyes at least) no value in fragmenting instances of the same character.

But no-one is fragmenting instances of the same character. They _are_ different characters from different languages. To take an example from the article, I am not sure how I feel about combining B and β. You are either ignoring the whole of English speaker population or the greek speaking one. Given that you have complete flexibility to assign a code for both of them, why not do it(responsibly)?

Re: I Can’t Write My Name in Unicode

#55

Wait, "ত + ্ + ‍ = ‍ৎ" is nothing like "\ + / + \ + / = W". The Bengali script is (mostly) an abugida. Ie, consonants have an inherent vowel (/ɔ/ in the case of Bengali), which can be overriden with a diacritic representing a different vowel. To write /t/ in Bengali, you combine the character for /tɔ/, "ত", with the "vowel silencing diacritic" to remove the /ɔ/, " ্". As it happens, for "ত", the addition of the diacr…

I don't understand Bengali at all, but the character ৎ does have it's own Unicode codepoint (U+09CE, BENGALI LETTER KHANDA TA). It was introduced in Unicode 4.1 in 2005.

Re: I Can’t Write My Name in Unicode

#57

I don't understand; I don't feel like character combination using the zero width joiner is on the same level as 13375p34k. It looks like the character just doesn't have a separate code-point, but is instead a composite, but still technically "reachable" from within Unicode, no?

It's like typing ` + o to get "ò", isn't it? You can argue that ò is actually an o with that tilde, while that character is not ত + ্ + an invisible joining character, but that's an input method thing, and there is a ৎ character after all.

They're on the same key on my keyboard, but ` is a grave, ~ is a tilde.

Re: I Can’t Write My Name in Unicode

#58
post #45
post #11

Earlier quoted context omitted.

Bengali is the seventh most widely spoken language in the world, with more native speakers than Russian. That's hardly a "niche language". And, as pg has pointed out, making the computer industry English-centric and America-centric is greatly limiting us. Intelligence and will to power are equally distributed across the globe. But institutional barriers to success are not. This is a classic example of an institutiona…

Right, but some languages are insanely complex to implement. It might be a better idea to teach English to people around the globe rather than cater to every individual need (which will still leave people unable to communicate across languages). I'm not saying other languages should go away -- but the world would also benefit from having a "universal" language, which is more or less English at this point (Mandarin is…

> some languages are insanely complex to implement

I don't understand this, is there more to implementing a language than creating glyphs for its character set? I wouldn't think the linguistic complexity would matter at all, only the number of glyphs in the 'alphabet' or similar?

Re: I Can’t Write My Name in Unicode

#59

Wait, "ত + ্ + ‍ = ‍ৎ" is nothing like "\ + / + \ + / = W". The Bengali script is (mostly) an abugida. Ie, consonants have an inherent vowel (/ɔ/ in the case of Bengali), which can be overriden with a diacritic representing a different vowel. To write /t/ in Bengali, you combine the character for /tɔ/, "ত", with the "vowel silencing diacritic" to remove the /ɔ/, " ্". As it happens, for "ত", the addition of the diacr…

Native Bengali here. The ligature used for the last letter of "haTaat" (suddenly) is not the same as the last ligature in "aditya" - the latter doesn't have the circle at the top.

More generally, using the vowel silencing diacritic (hasanta) along with a separate ligature for the vowel ending - while theoretically correct - does not work because no one writes that way! Not using the proper ligatures makes the test essentially unreadable.

Re: I Can’t Write My Name in Unicode

#60
post #13
post #4

I wonder if the author has submitted a proposal to get the missing glyph for their name added. You don't need to be a member of the consortium to propose adding a missing glyph/updating the standard. The point of the committee as I understand it isn't to be an expert in all forms of writing, but to take the recommendations from scholars/experts and get a working implementation, though more diverse representation of l…

The author's explanation of what characters Chinese, Japanese, and Korean share is very limited. All three languages use Chinese characters in written language to varying extents, and in some cases the differences begin significantly less than a century ago. Though there are cases where the same Chinese character represented in Japanese writing is different from how it is represented in Traditional Chinese writing (i…

Also Antiqua and Fraktur used to be seen as different writing systems (ſ, I and J being equivalent, tironian et being some examples where they differ), yet this is largely ignored by Unicode (except when used in mathematics)
Post reply on HN