Live data from Hacker News

I Can’t Write My Name in Unicode

modelviewculture.com

21–30 of 377 posts

Re: I Can’t Write My Name in Unicode

#21
post #10
post #5

I get the author's point, but assigning blame to the Unicode Consortium is incorrect. The Indian government dropped the ball here. They went their separate way with "ISCII" and other such misguided efforts, instead of cooperating with UC. To me, the UC is just a platform. The government is the de facto safeguarder of the people's interests; if it drops the ball, it should be taken to task, not the provider of the pla…

This is a terrible excuse: the Unicode Consortium should always seek out at least one (if not a group of) native speakers of a language before defining code points for that language. These speakers really should be both native speakers of and experts in the language. There are countless ways to reach out to Bengali speakers, only one of which is the Indian government - whatever politics governments may play, a techno…

That's ridiculous. How can you expect UC to reach out to the 1000s of languages out there? Plus, shouldn't the party who expects to benefit put in the effort?

Did you know that the Indian Government has a department for just such a thing, http://tdil.mit.gov.in/ ? What has it been doing all these years?

Added later: there's also CDAC http://cdac.in/index.aspx?id=mlingual . And these are just 2 that I found quickly.

Re: I Can’t Write My Name in Unicode

#22
post #3

By the title, I assumed this would be an article about how 2-year-olds can learn to do some interesting stuff on a smartphone. [Edit -- Sorry about this lame joke. I've suffered the consequences in downvotes.]

I thought it was going to be about illiteracy.

Re: I Can’t Write My Name in Unicode

#23
post #5

I get the author's point, but assigning blame to the Unicode Consortium is incorrect. The Indian government dropped the ball here. They went their separate way with "ISCII" and other such misguided efforts, instead of cooperating with UC. To me, the UC is just a platform. The government is the de facto safeguarder of the people's interests; if it drops the ball, it should be taken to task, not the provider of the pla…

But then consider the implementation path to fixing the problem for a minority linguistic group being deliberately repressed by their government--It would require blood. If there is some alternate process to work with the UC directly, that could be better but it puts the UC in the position of judging a linguistic group's claims for legitimacy.

I agree that this refutes claims that the UC was negligent, but we can still say that they failed to be particularly assiduous in this case.

Side question: does anyone know the story that led to http://unicode.org/charts/PDF/U13A0.pdf ?

EDIT: subjunctive

Re: I Can’t Write My Name in Unicode

#24
post #5

I get the author's point, but assigning blame to the Unicode Consortium is incorrect. The Indian government dropped the ball here. They went their separate way with "ISCII" and other such misguided efforts, instead of cooperating with UC. To me, the UC is just a platform. The government is the de facto safeguarder of the people's interests; if it drops the ball, it should be taken to task, not the provider of the pla…

There are over 100 languages spoken in India, many of which are not even Hindustani in origin. The Indian government primarily recognizes two languages: Hindi and Urdu. Additionally, Urdu is the official national language of Pakistan. Are speakers of Bengali, Tamil, Marathi, Punjabi, etc. really supposed to depend on the Indian government to assure that the Unicode Consortium supports their native tongues? What about…

The Indian government recognizes 22 languages, not 2. The government has a department dedicated to their support: http://tdil.mit.gov.in Why isn't someone asking that department WTF has it been doing?

Added later: and there's also CDAC http://cdac.in/index.aspx?id=mlingual which has been at it since 1988.

Re: I Can’t Write My Name in Unicode

#25
post #13
post #4

I wonder if the author has submitted a proposal to get the missing glyph for their name added. You don't need to be a member of the consortium to propose adding a missing glyph/updating the standard. The point of the committee as I understand it isn't to be an expert in all forms of writing, but to take the recommendations from scholars/experts and get a working implementation, though more diverse representation of l…

The author's explanation of what characters Chinese, Japanese, and Korean share is very limited. All three languages use Chinese characters in written language to varying extents, and in some cases the differences begin significantly less than a century ago. Though there are cases where the same Chinese character represented in Japanese writing is different from how it is represented in Traditional Chinese writing (i…

Nor is it unreasonable to "unify" Latin, Greek and Cyrilic:

Cyrillic ПФ vs Greek ΠΦ

Cyrillic АВ vs Latin AB

Obviously using ω for w (as he does) is stupid, but his reducto-ad-absurdum is not particularly absurd.

Re: I Can’t Write My Name in Unicode

#26
post #4

I wonder if the author has submitted a proposal to get the missing glyph for their name added. You don't need to be a member of the consortium to propose adding a missing glyph/updating the standard. The point of the committee as I understand it isn't to be an expert in all forms of writing, but to take the recommendations from scholars/experts and get a working implementation, though more diverse representation of l…

I found it extremely annoying that he doesn't even say specifically what the problem with his name is . So there's a letter that's unavailable? Which letter?

> Even today, I am forced to do this when writing my own name. My name is not only a common Indian name, but one of the top 1,000 names in the United States as well. But the final letter has still not been given its own Unicode character, so I have to use a substitute.

Not as descriptive as it could be, but this article isn't about him.

Re: I Can’t Write My Name in Unicode

#27
post #17

> He proudly announces that there are ‘no fewer than 147 Indian dialects’ – a pathetically inaccurate count. (Today, India has 57 non-endangered and 172 endangered languages, each with multiple dialects – not even counting the many more that have died out in the century since My Fair Lady took place) So, how many were there really? At the time, I mean.

I think the numbers are somewhat disputed. The People's Linguistic Survey of India says there are at least 780, with ~220 having died out in the last half century.[1] The Anthropological Survey of India reported 325 languages.[2]

The discrepancies are made particularly tricky because of the somewhat ambiguous distinction between languages and dialects.

[1] http://blogs.reuters.com/india/2013/09/07/india-speaks-780-l...

[2] http://books.google.com/books?id=VjGdDo75UssC&pg=PA145

Re: I Can’t Write My Name in Unicode

#28

I don't understand; I don't feel like character combination using the zero width joiner is on the same level as 13375p34k. It looks like the character just doesn't have a separate code-point, but is instead a composite, but still technically "reachable" from within Unicode, no?

It's like typing ` + o to get "ò", isn't it? You can argue that ò is actually an o with that tilde, while that character is not ত + ্ + an invisible joining character, but that's an input method thing, and there is a ৎ character after all.

Re: I Can’t Write My Name in Unicode

#29

Earlier quoted context omitted.

It sounds like the glyph is in unicode already, but expressed using combining characters?

Imagine if the letter Q had been left out of Unicode's Latin alphabet. The argument against it is that it can be written with a capital O combined with a comma. (That's going to play hell with naive sorting algorithms, of course, but oh well.) Oh, and also imagine your name is Quentin.

Combining characters are not just used for Bengali though. E.g. umlauted letters in European languages can also be expressed using combining characters, and implementations need to deal with those when sorting.

Re: I Can’t Write My Name in Unicode

#30
Wait, "ত + ্ + ‍ = ‍ৎ" is nothing like "\ + / + \ + / = W".

The Bengali script is (mostly) an abugida. Ie, consonants have an inherent vowel (/ɔ/ in the case of Bengali), which can be overriden with a diacritic representing a different vowel. To write /t/ in Bengali, you combine the character for /tɔ/, "ত", with the "vowel silencing diacritic" to remove the /ɔ/, " ্". As it happens, for "ত", the addition of the diacritic changes the shape considerably more than it usually does, but it's a perfectly legitimate to suggest the resulting character is still a composition of the other two (for a more typical composition, see "ঢ "+" ্" = "ঢ্").

As it happens, the same character ("ৎ") is also used for /tjɔ/ as in the "tya" of "aditya". Which suggests having a dedicated code point for the character could make sense. But Unicode isn't being completely nutso here.

Post reply on HN