Live data from Hacker News

I Can’t Write My Name in Unicode

modelviewculture.com

241–250 of 377 posts

Re: I Can’t Write My Name in Unicode

#241

Earlier quoted context omitted.

Imagine if the letter Q had been left out of Unicode's Latin alphabet. The argument against it is that it can be written with a capital O combined with a comma. (That's going to play hell with naive sorting algorithms, of course, but oh well.) Oh, and also imagine your name is Quentin.

My first choice as theoretical Quentin wouldn't be "how can I frame this accidental, perhaps even flagrantly disrespectful omission as antiprogressive and dissect the credentials, experience, and ethnicity of the people who made the mistake via culture essay," it would probably be "where do I issue a pull request to fix this mistake or in what way can I help?" Maybe that's just me. I look forward to the future where…

Anyone can propose the addition of a new character to Unicode. It doesn't take $18,000 as some people think. You just need to convince the Unicode Consortium that it makes sense (preferably with solid evidence on use of the character). The process is discussed at: http://unicode.org/pending/proposals.html

I have a proposal of my own in the works to add a character to Unicode, so I'll see how it goes. There's a discussion of how someone successfully got the power symbol added to Unicode at https://github.com/jloughry/Unicode so take a look if you're thinking of proposing a character.

Re: I Can’t Write My Name in Unicode

#242
post #13
post #4

I wonder if the author has submitted a proposal to get the missing glyph for their name added. You don't need to be a member of the consortium to propose adding a missing glyph/updating the standard. The point of the committee as I understand it isn't to be an expert in all forms of writing, but to take the recommendations from scholars/experts and get a working implementation, though more diverse representation of l…

The author's explanation of what characters Chinese, Japanese, and Korean share is very limited. All three languages use Chinese characters in written language to varying extents, and in some cases the differences begin significantly less than a century ago. Though there are cases where the same Chinese character represented in Japanese writing is different from how it is represented in Traditional Chinese writing (i…

Exactly. The Han Unificiation project never tried to unify everything. They just took the set of common characters and unified them, leaving the rest alone. They may have made some mistakes in choosing which characters to unify, but for the most part they did a splendid job.

Korean (Hangul) has its own, massive block of over 11K code points. Japanese (Hiragana, Katakana, and other assorted symbols) also has its own block outside of the Unified Han block. Chinese characters that are clearly distinct get their own code points as well. How else would I write 国 and 國 in the same sentence?

Re: I Can’t Write My Name in Unicode

#243
post #216

Earlier quoted context omitted.

Was British rule actually a net negative, in retrospect? Have there been studies done using objective criteria (not emotional) over counties that were colonies versus ones that weren't? I suppose you can't really quantify the value of people that were destroyed by colonization, but you can look at the current population. Also I just gotta wonder: suppose European or other relatively simple-to-encode languages didn't…

>" Was British rule actually a net negative, in retrospect? Have there been studies done using objective criteria (not emotional) over counties that were colonies versus ones that weren't? " It's both a positive, and a negative. But if you want to calculate whether it was a "net" positive or "net" negative, then I'm afraid you're going to have to quantify very difficult-to-quantify concepts. And you'll have to do it…

Well that's why I said without taking into account people that were destroyed by it. That is, if we look purely at modern-day circumstances, are they better?

A comparison might be a country with overpopulation. If you kill x% of the population and enact some birth control plan, then 100 years later you might end up with a "net positive" discounting the people that were killed and the emotional effects of their loved ones.

Are there countries in Africa that weren't colonized, or some islands that weren't, that we could compare to ones that were and make some inferences?

Re: I Can’t Write My Name in Unicode

#244
post #188

Earlier quoted context omitted.

> > It seems like you're trying to single out combining pairs as "less legitimate" when they're extensively used in the standard. > I'm saying that Unicode only does it in English where it makes semantic sense to a native English speaker. Well, combining characters almost never come up in English. The best I can think of would be the use of cedillas, diaereses, and acute accents in words like façade, coördinate and…

> ch, ll, ñ, and rr are all considered separate letters In Spanish "rr" has never been considered as a single letter. "Ch" and "ll" used to be, but not anymore. Ñ is, of course.

That's funny; I spent several hours of class time trilling r's to make sure we pronounced "carro" correctly, and repeating a 30 character alphabet.

Re: I Can’t Write My Name in Unicode

#245
post #48

Not sure if the l33tspeak analogy is fully justified. In case of the "missing" letter (called khanda-ta in Bengali) for the Bengali equivalent of "suddenly", historically, it has been a derivative of the ta-halant form (ত + ্ + ‍ ). As the language evolved, khanda-ta became a grapheme of its own, and Unicode 4.1 did encode it as a distinct grapheme. A nicely written review of the discussions around the addition can b…

> I could write the author's name fine: আদিত্য Author here. Well, yes and no. The jophola at the end is not actually given its own codepoint[0]. The best analogy I can give is to a ligature in English[1]. The Bengali fonts that you have installed happen to render it as a jophola, the way some fonts happen to render "ff" as "ff" but that's not the same thing as saying that it actually is a jophola (according to the Uni…

Doesn't this reply invalidate your whole point of the article?

Seems like there was a lot of hard work put in on making the "khanda-ta" work properly?

Re: I Can’t Write My Name in Unicode

#246
post #211

Earlier quoted context omitted.

> > It seems like you're trying to single out combining pairs as "less legitimate" when they're extensively used in the standard. > I'm saying that Unicode only does it in English where it makes semantic sense to a native English speaker. Well, combining characters almost never come up in English. The best I can think of would be the use of cedillas, diaereses, and acute accents in words like façade, coördinate and…

> Portuguese, on the other hand, doesn't officially include k > or y in the alphabet. With no judgement towards your broader point, I'd like to point out that this is no longer the case as of the orthographic agreement of 1990[0]. As far as I know it's been added back in order to better suit African speakers. [0] https://pt.wikipedia.org/wiki/Acordo_Ortogr%C3%A1fico_de_199...

That's good to know. I learned Portuguese in '97-'99, so the information I had was incorrect at the time. We Anericans always recited the alphabet with k and y, but our teacher said they weren't official (although he also said that Brazilians would recognize them).

Re: I Can’t Write My Name in Unicode

#247
post #239
post #221

Earlier quoted context omitted.

Except you can still recognize the 'a' as 'a' no matter which way it is rendered. Not so with Chinese characters. For instance, the character for "fly" in simplified (飞) and traditional (飛) look very different. Someone who only learned simplified may not recognize the traditional character as being the same.

Which is exactly why 飞 and 飛 are encoded separately. I don't see any problem with that.

Yes, but other characters that also look different are merged. Here's an example: http://www.tofugu.com/2012/04/04/the-sorry-state-of-japanese...

That's the character for "cold". If you showed me (a Chinese speaker) the Japanese or Korean variant, I would have no idea what it meant.

Re: I Can’t Write My Name in Unicode

#248
I see many comments about Han unification being a bad idea but I am not seeing any reason why it was such a bad idea. I am from a CJK country and I find it makes a lot of sense. Most commonly used characters should be considered identical regardless of whether it is used in Chinese Japanese Korean or Vietnamese. Sure there are some characters that are rendered improperly depending on your font but I don't think that makes Han unification a fundamentally bad idea.

Re: I Can’t Write My Name in Unicode

#249
post #227
post #156

It sounds like the author is looking to be offended. Talking about Bengali being the seventh largest native language, then saying that US$18k is too expensive for a stake in solving this problem? Emojis with skin tones, something every human can use, that have arrived before a character that only Bengalis use is taken as an " outright insult "?

> Talking about Bengali being the seventh largest native language, then saying that US$18k is too expensive for a stake in solving this problem? West Bengal and Bangladesh aren't exactly the richest places in the world. > Emojis with skin tones, something every human can use, that have arrived before a character that only Bengalis use is taken as an "outright insult"? Nobody is being inconvenienced by the inability t…

$18k is really a small amount of money for any governmental organization including those places. Even North Korea participates in the process.

Re: I Can’t Write My Name in Unicode

#250
Most of the discussion here centers on how linguistic commitments ought to drive decision making in determining the Unicode spec. It's already been covered how Unicode provides a number to symbol (i.e. codepoint) mapping; composition, rendering, input, etc. is left to the system implementation to determine.

I actually worked with Lee Collins (he was my manager) and a whole bunch of the ITG crew at Apple in the early 00's. The critique that this was just a bunch of white dudes that only loved their own mother tongues but studied other languages dispassionately as a superficial proof of their right to determine an international encoding spec is, to me, misinformed, and DEEPLY offensive.

I had just gotten out of college and it was AMAZING to see how much these people loved language! Like it was almost weird. Our department had a very diverse linguistic background and, oftentimes, when you had a question about a specific language property you could just go down the hall and just ask "Is this the way it should be?"

All the discussion here happened on the unicode mailing lists as well as in passing at lunch and at team meetings. Lots of people felt VERY passionately about particular things; I just liked watching the drama. But to write an article like this that somehow intimates that people didn't care is wrong. People cared A LOT.

It's been touched in a couple of comments, but a big factor in Unicode was also the commercial viability of certain decisions. You have Apple, Adobe, Microsoft, etc with all of their previous attempts at internationalization/localization. If you wanted these large companies to adopt a new standard you had to give them a relative easy pathway to conversion and adoption.

I think the article, in general, lacks a perspective that dismisses the work of the Unicode team as well as the different stakeholders. Historically, accessible computing was birthed in the United States. The tools and processes are naturally Ameri-centric. I'm not saying it's right, I'm just saying it is. The fact that the original contributors to Unicode were (mostly?) American shouldn't be a surprise; they had the most at stake in creating interoperable systems that spanned linguistic boundaries.

http://www.unicode.org/history/earlyyears.html

A purely academic solution may have been better. That's not certain. But it's pretty clear that any solution that didn't address the needs of pre-existing commercial systems would never be adopted. I'm surprised this hasn't been emphasized more.

I have more thoughts on how to solve the problem that the author complains about but I'll leave that for another day.

Post reply on HN