Live data from Hacker News

I Can’t Write My Name in Unicode

modelviewculture.com

171–180 of 377 posts

Re: I Can’t Write My Name in Unicode

#171
post #48

Not sure if the l33tspeak analogy is fully justified. In case of the "missing" letter (called khanda-ta in Bengali) for the Bengali equivalent of "suddenly", historically, it has been a derivative of the ta-halant form (ত + ্ + ‍ ). As the language evolved, khanda-ta became a grapheme of its own, and Unicode 4.1 did encode it as a distinct grapheme. A nicely written review of the discussions around the addition can b…

> I could write the author's name fine: আদিত্য Author here. Well, yes and no. The jophola at the end is not actually given its own codepoint[0]. The best analogy I can give is to a ligature in English[1]. The Bengali fonts that you have installed happen to render it as a jophola, the way some fonts happen to render "ff" as "ff" but that's not the same thing as saying that it actually is a jophola (according to the Uni…

[deleted]

Re: I Can’t Write My Name in Unicode

#172
post #169
post #147

Earlier quoted context omitted.

Absolutely. But the two "g" haven't diverged, yet. We don't give out code points to speculative future developments. If and when they diverge one will get its very own code point.

The premise of Someone's hypothetical was to explain to theon144 why it might be both hard and important to guess. The hypothetical assumed that the difference already existed. It echos the larger context of Han unification that com2kid started, but with an example that's a lot easier for native English speakers to understand. So while I agree that they haven't yet diverged, that's outside the context of the hypothet…

That's wrong. Someone's hypothetical had no diverging meaning involved, only a stylistic choice in the presentation of the same letter.

And he is still wrong when it comes to the claimed necessity to "guess".

The user has set up his system correctly. "Guessing" only comes into play if you want to force your stylistic variation on others.

And that is obviously a bad idea, for all the reasons called out before, like familiarity and readability.

Let's not kid ourselves. When it comes to Han unification opposition there is mostly one issue at play: plain racism. "Our holy script shall not be defiled by those dirty bastards". And that works in all directions.

Re: I Can’t Write My Name in Unicode

#173
post #5

I get the author's point, but assigning blame to the Unicode Consortium is incorrect. The Indian government dropped the ball here. They went their separate way with "ISCII" and other such misguided efforts, instead of cooperating with UC. To me, the UC is just a platform. The government is the de facto safeguarder of the people's interests; if it drops the ball, it should be taken to task, not the provider of the pla…

But then consider the implementation path to fixing the problem for a minority linguistic group being deliberately repressed by their government--It would require blood. If there is some alternate process to work with the UC directly, that could be better but it puts the UC in the position of judging a linguistic group's claims for legitimacy. I agree that this refutes claims that the UC was negligent, but we can sti…

In specific or in general? http://en.wikipedia.org/wiki/Cherokee_syllabary

An increasing corpus of children's literature is printed in Cherokee syllabary today to meet the needs of Cherokee students in the Cherokee language immersion schools in Oklahoma and North Carolina. In 2010, a Cherokee keyboard cover was developed by Roy Boney, Jr. and Joseph Erb, facilitating more rapid typing in Cherokee and now used by students in the Cherokee Nation Immersion School, where all coursework is written in syllabary.[8] The syllabary is finding increasingly diverse usage today, from books, newspapers, and websites to the street signs of Tahlequah, Oklahoma and Cherokee, North Carolina.

Re: I Can’t Write My Name in Unicode

#174
post #170

Earlier quoted context omitted.

Not to say there aren’t problems with CJK unification. It’s difficult or impossible to represent old family names or newly coined characters without some means of composing characters from radicals—even if you do want “precomposed” characters a majority of the time, as is the case with, say, é (00E9) versus é (0065 0301). As far as these characters being “the same”, I think it’s better to say they’re analogues . A i…

> I would actually support the author’s “Greco Unification” strawman if it could be done in a principled way A 5 year old tech note by a Unicode Consortium member at http://www.unicode.org/notes/tn26 gives 7 reasons "why the Latin, Greek, and Cyrillic scripts have been separately encoded, rather than being encoded as a single script", and 12 reasons why the Han script was unified. Edit: These 7 reasons make a good ca…

[deleted]

Re: I Can’t Write My Name in Unicode

#175
post #5

I get the author's point, but assigning blame to the Unicode Consortium is incorrect. The Indian government dropped the ball here. They went their separate way with "ISCII" and other such misguided efforts, instead of cooperating with UC. To me, the UC is just a platform. The government is the de facto safeguarder of the people's interests; if it drops the ball, it should be taken to task, not the provider of the pla…

> The Indian government dropped the ball here. They went their separate way with "ISCII" and other such misguided efforts, instead of cooperating with UC

According to the beginning of chapter 12 of the Unicode Standard 7.0:

"The major official scripts of India proper [...] are all encoded according to a common plan, so that comparable characters are in the same order and relative location. This structural arrangement, which facilitates transliteration to some degree, is based on the Indian national standard (ISCII) encoding for these scripts. The first six columns in each script are isomorphic with the ISCII-1988 encoding, except that the last 11 positions, which are unassigned or undefined in ISCII-1988, are used in the Unicode encoding."

Re: I Can’t Write My Name in Unicode

#176
post #70

Earlier quoted context omitted.

Han unification has been overly aggressive about merging some characters, but the basic principle is not as flawed as it is some times (as in this article) made to sound. The vast majority of Japanese and Chinese characters are not only similar, they are identical. Not all are. Some are clearly different characters deriving from a common historical root, and should not be unified. Sometimes, when characters are a bit…

> You don't want a in German and a in English to be different letters just because Helvetica and Baskerville look different. A much more apt comparison would be Antiqua and Fraktur. It was a commonly-held belief that Fraktur was the authentic German alphabet, separate from the alphabets used by other languages. Back in the 19th century, English-German dictionaries even used Antiqua for English words and Fraktur for G…

no, it isn't a more apt comparison. Antiqua and Fraktur look much more different from each other than most Chinese and Japanese characters are from eachother.

Here are two screenshots in Japanese and Chinese, taken from the http://www.yomiuri.co.jp/ for one, and http://www.cntv.cn/ for the other. Neither media outlet can be accused of being cultural sellouts misrepresenting the written tradition of their countries.

Japanese: https://www.dropbox.com/s/lpb3dk26ftds6oh/Screenshot%202015-... Chinese: https://www.dropbox.com/s/dcwv33ho0juracg/Screenshot%202015-...

These are the same characters, and unifying them is the only sane thing to do. The same is true for most characters. There are also some characters which are clearly distinct in both scripts, and unicode treats them as different. Some are grey-zone, and there is room for reasonable disagreement.

But as a whole, the case of unification of Chinese and Japanese is stronger than for Antiqua and Fraktur.

Re: I Can’t Write My Name in Unicode

#177

I am an Indian and it shocks me that Indians are still blaming the British after 70 yrs of independence. Is 70 years of Independence not enough to make your language "first class citizen" ? Ofcourse Bengali is second class language because Bengalis didn't invent the standard. Can we stop blaming white people for everything. Seriously WTF.

Was British rule actually a net negative, in retrospect? Have there been studies done using objective criteria (not emotional) over counties that were colonies versus ones that weren't? I suppose you can't really quantify the value of people that were destroyed by colonization, but you can look at the current population. Also I just gotta wonder: suppose European or other relatively simple-to-encode languages didn't…

It is really hard to quantify something like that. While British siphoned off large amounts of natural resources and often caused widespread famines in India. They also gave us western democracy, education in western science,access to all the western knowledge.

Re: I Can’t Write My Name in Unicode

#178
post #17

> He proudly announces that there are ‘no fewer than 147 Indian dialects’ – a pathetically inaccurate count. (Today, India has 57 non-endangered and 172 endangered languages, each with multiple dialects – not even counting the many more that have died out in the century since My Fair Lady took place) So, how many were there really? At the time, I mean.

I believe the number of "dialects" named in My Fair Lady can be largely explained by the lack of clear distinction between language and dialect over the years. From [1]: "There is no universally accepted criterion for distinguishing a language from a dialect. A number of rough measures exist, sometimes leading to contradictory results. The distinction is therefore subjective and depends on the user's frame of reference."

Getting upset about Henry Higgins's estimation of the number of Indian "dialects" in a play from many decades ago doesn't make sense to me. His character was deliberately portrayed as a regressive lout, and terminology has surely changed in the intervening years.

[1]: http://en.wikipedia.org/wiki/Dialect#Dialect_or_language

Re: I Can’t Write My Name in Unicode

#179
post #139
post #82

Earlier quoted context omitted.

I'm fluent in Japanese and speak some Mandarin Chinese as well. These 3 characters are identical, not similar. For a different example, 国 and 國 used to be the same character, but China and Japan (left) have both diverged the traditional form still used in Taiwan (right). Unicode treats them as separate. 今 Looks slightly different in traditional Chinese vs other languages. In traditional Chinese, the little straight l…

This seems again to be a perfect place for rendering rather than encoding. The english letter 'a' can be rendered as a ring with a tail (the way I handwrite), or a ring with a cap and a tail (the way the font usually renders). Both are the same letter, if rendered differently based on my (contextually sensitive) font.

I think not. I do want to be able to say both People's Republic of China 中华人民共和国 and Republic of China 中華民國 in the same text, and if I had to choose rendering either 国+华 or 國+華 then it wouldn't work.

Re: I Can’t Write My Name in Unicode

#180
post #129

Earlier quoted context omitted.

Actually, I can't. I can decipher quite a bit with lots of effort, but I wouldn't call that "reading". And there are very few people around who can read all those (especially Textura) without problems. On your other example: when I was in Russia, I found those "unknown" letters difficult, but fun. But the letters they share with Latin? I didn't see a difference. And to your phone example: of course, I'd be annoyed. B…

> On your other example: when I was in Russia, I found those "unknown" letters difficult, but fun. But the letters they share with Latin? I didn't see a difference. Again, Han unification does this. Most characters are the same, great! But some are different. Sucks for those that are different. > And to your phone example: of course, I'd be annoyed. But that's exactly the point: you have to set up your local system c…

As far as I know, those "some are different" characters have not been unified, but got separate code points. And that's where your whole argumentation falls apart, unless you claim that the Unicode consortium mis-classified lots of characters.

But then I'd just retreat, because I don't speak any Asian language and cannot verify the claim myself. I can only defer to the experts, and they say that issue has been taken care of.

As far as space savings are concerned, I replied to that in another comment to you. It's not okay to dismiss Han unification as just some space saving sttempt that's not needed anymore. Space saving was one motivation, but according to the consortium not the primary one.

Post reply on HN