Live data from Hacker News

I Can’t Write My Name in Unicode

modelviewculture.com

101–110 of 377 posts

Re: I Can’t Write My Name in Unicode

#101
post #87
post #70

Earlier quoted context omitted.

Han unification has been overly aggressive about merging some characters, but the basic principle is not as flawed as it is some times (as in this article) made to sound. The vast majority of Japanese and Chinese characters are not only similar, they are identical. Not all are. Some are clearly different characters deriving from a common historical root, and should not be unified. Sometimes, when characters are a bit…

> The vast majority of Japanese and Chinese characters are not only similar, they are identical. Not all are. Some are clearly different characters deriving from a common historical root, and should not be unified. And what about traditional versus simplified? Which glyph set do I use? Oh wait, thanks to Han unification, I now need to rely on bloody environment variables to decide! For Chinese text, rendering a strin…

Antiqua, Fractura, Schwabacher, Textura and all those other charming variants of writing European languages don't have seperate code points for all those presentational variations.

Do we white Western men happen to discriminate against ourselves?

Re: I Can’t Write My Name in Unicode

#102
post #85

Earlier quoted context omitted.

do you have any evidence that the UC has actively ignored requests from Bengali speakers? Has any Bengali speaker made proposals to the UC for fixing these issues? If yes, and the UC chose to ignore them, then there is some blame to be assigned with the UC. Otherwise, this is a non-issue. Take, for example, Tibetan. The number of Tibetan speakers is minuscule compared to, say, Bengali. But still Tibetan has good supp…

[deleted]

Not true. Anybody with an email address can send feedback. And individuals who want to be more active and become an actual member have a $75 membership fee.

Re: I Can’t Write My Name in Unicode

#103
post #45

Earlier quoted context omitted.

Right, but some languages are insanely complex to implement. It might be a better idea to teach English to people around the globe rather than cater to every individual need (which will still leave people unable to communicate across languages). I'm not saying other languages should go away -- but the world would also benefit from having a "universal" language, which is more or less English at this point (Mandarin is…

> some languages are insanely complex to implement I don't understand this, is there more to implementing a language than creating glyphs for its character set? I wouldn't think the linguistic complexity would matter at all, only the number of glyphs in the 'alphabet' or similar?

It really depends on the language, and it's not totally about the glyphs. Text entry is a huge challenge for languages like Mandarin where glyphs can have multiple pronunciations and meanings depending on context. Consider that Mandarin (which shares many, but not all glyphs with Japanese kanji) has upwards of 20,000 different glyphs, and that other languages have a similar level of complexity, and it becomes hard to find an encoding standard capable of handling all of that complexity and variance.

What constitutes a "glyph" isn't even consistent - in some languages a glyph is a syllable, in some (like English) it's less than a syllable, and in yet others a single glyph can be an entire word.

In a language like Japanese, multiple glyphs are often combined to create new composite glyphs with different meanings. For example, the word for "forest" is a glyph comprised of 3 "tree" glyphs, but has an unrelated pronunciation.

How do you handle text entry between these differences? It may seem like a pedantic question, but it makes sense to define the characters in the way they will be written, or else the text entry scheme will be so complex you'll need an interpreter to convert from some entry scheme into the Unicode format. I think this is the problem the Unicode Consortium is grappling with - and it's not an easy problem. I don't claim to have the answers here; but I do recognize the complexity.

Re: I Can’t Write My Name in Unicode

#104
post #73

Earlier quoted context omitted.

Isn't the idea technically that the code shouldn't even have to guess? Why isn't this the case?

Imagine a world where the British always write the lowercase letter g as a single-story glyph ( http://en.m.wikipedia.org/wiki/G#Typographic_variants ). The colonies start writing it identically, but after a while, they start writing it as a double-story g. After a century or so, nobody in he colonies writes the single-story variant, and all Brits always do. The unicode consortium studies the case and concludes that…

No, you don't. It's the same letter and it's always "goto".

Your local setup determines the look of the glyph, so nobody sees an unfamiliar form.

But maybe you'd like to encode typefaces/fonts in the Unicode code points, as well? To make sure that I'm seeing the exact same arrangement of pixels you want me to see?

Re: I Can’t Write My Name in Unicode

#105
post #89
post #81

Earlier quoted context omitted.

You don't determine language based on codepage. I give you ASCII text; what language is it?

BINGO. But because of Han unification I all of a sudden DO need to know the language. The same Unicode code point needs to be rendered differently for a user in Mainland China versus a user in Japan or else the user may not be able to read the text! Even if the user can read the character, they are going to experience a degradation in reading speed and comprehension, and be generally frustrated. Not to mention showin…

Sorry, Han glyphs render the same in Chinese and Japanese.

Regarding simplified versus traditional, no one is seriously unifying those.

There's some minor disagreements as to when a minor stylistic or historical variant deserves a separate glyph, but this isn't about rendering different glyphs in Chinese or Japanese. If Unicode is doing its job no one should have difficulty reading unified Han characters in one font regardless of language.

Re: I Can’t Write My Name in Unicode

#106
post #48

Not sure if the l33tspeak analogy is fully justified. In case of the "missing" letter (called khanda-ta in Bengali) for the Bengali equivalent of "suddenly", historically, it has been a derivative of the ta-halant form (ত + ্ + ‍ ). As the language evolved, khanda-ta became a grapheme of its own, and Unicode 4.1 did encode it as a distinct grapheme. A nicely written review of the discussions around the addition can b…

> I could write the author's name fine: আদিত্য Author here. Well, yes and no. The jophola at the end is not actually given its own codepoint[0]. The best analogy I can give is to a ligature in English[1]. The Bengali fonts that you have installed happen to render it as a jophola, the way some fonts happen to render "ff" as "ff" but that's not the same thing as saying that it actually is a jophola (according to the Uni…

Unicode makes extensive use of combining characters for european languages, for example to produce diacritics: ìǒ or even for flag emoji. A correct rendering system will properly combine those, and if it doesn't then that's a flaw in the implementation, not the standard. It seems like you're trying to single out combining pairs as "less legitimate" when they're extensively used in the standard.

Re: I Can’t Write My Name in Unicode

#107
post #16
post #4

I wonder if the author has submitted a proposal to get the missing glyph for their name added. You don't need to be a member of the consortium to propose adding a missing glyph/updating the standard. The point of the committee as I understand it isn't to be an expert in all forms of writing, but to take the recommendations from scholars/experts and get a working implementation, though more diverse representation of l…

Holy cow CJK unification is a terrible idea. Maybe if it originated from the CJK governments, it might be an OK idea, but the idea of a Western multinationals trying to save Unicode space by disregarding the distinctness of a whole language group is idiotic. The fundamental roll of an institution like the Unicode Consortium is to be descriptive, not prescriptive. If there is a human script, passing certain, low, low…

To oppose Han unification is to say that an 'a' in English and an 'a' in French should be different code points, because they're different languages.

Alternatively, if Unicode directly encoded words, rather than letters, of Western languages akin to ideographs in East Asian languages, it's like arguing that 'color' and 'colour' should be separate code points.

Re: I Can’t Write My Name in Unicode

#108

Earlier quoted context omitted.

Imagine if the letter Q had been left out of Unicode's Latin alphabet. The argument against it is that it can be written with a capital O combined with a comma. (That's going to play hell with naive sorting algorithms, of course, but oh well.) Oh, and also imagine your name is Quentin.

> Imagine if the letter Q had been left out of Unicode's Latin alphabet. To properly write my european last name I have to press between 2 and 4 different simultaneous keys, depending on the system. Han unification is beyond misguided, but combining characters is not the problem.

Han unification as a hole is misguided? I'll grant you that some characters which were unified probably shouldn't have been, and maybe some that some that should have been weren't, but what's the argument for the whole thing to be misguided?

Should Norwegian A and English A be different Unicode code points just because Norwegian also has Ø, proving that it is a different writing system? You may want to debate whether i and ı should the same letter (they aren't), but most letters in the Turkish alphabet are the same as the letters in the English alphabet.

Re: I Can’t Write My Name in Unicode

#109
post #106

Earlier quoted context omitted.

> I could write the author's name fine: আদিত্য Author here. Well, yes and no. The jophola at the end is not actually given its own codepoint[0]. The best analogy I can give is to a ligature in English[1]. The Bengali fonts that you have installed happen to render it as a jophola, the way some fonts happen to render "ff" as "ff" but that's not the same thing as saying that it actually is a jophola (according to the Uni…

Unicode makes extensive use of combining characters for european languages, for example to produce diacritics: ìǒ or even for flag emoji. A correct rendering system will properly combine those, and if it doesn't then that's a flaw in the implementation, not the standard. It seems like you're trying to single out combining pairs as "less legitimate" when they're extensively used in the standard.

> Unicode makes extensive use of combining characters for european languages, for example to produce diacritics: ìǒ or even for flag emoji.

But it doesn't, for example say that a lowercase "b" is simply "a lowercase 'l' followed by an 'o' followed by an invisible joiner", because no native English speaker thinks of the character "b" as even remotely related to "lo" when reading and writing.

> It seems like you're trying to single out combining pairs as "less legitimate" when they're extensively used in the standard.

I'm saying that Unicode only does it in English where it makes semantic sense to a native English speaker. It does it in Bengali even where it makes little or no semantic sense to a native Bengali speaker.

Re: I Can’t Write My Name in Unicode

#110
post #48

Not sure if the l33tspeak analogy is fully justified. In case of the "missing" letter (called khanda-ta in Bengali) for the Bengali equivalent of "suddenly", historically, it has been a derivative of the ta-halant form (ত + ্ + ‍ ). As the language evolved, khanda-ta became a grapheme of its own, and Unicode 4.1 did encode it as a distinct grapheme. A nicely written review of the discussions around the addition can b…

> For me, many of these problems are more of an input issue, than an encoding issue. I think you've hit the nail on the head here. I'm a native English speaker, so I may in fact be making bad assumptions here, but I think the biggest issue here is that people conflate text input systems with text encoding systems. Unicode is all about representing written text in a way that computers can understand. But the way that…

They're not unrelated though. You have to have a way to get from your input format to the finished product in a consistent way, and the glyph set you design has a large bearing on that. You can't solve it completely with AI, because then you just have an AI interpretation of human language, not human language. A language like Korean written in Hangul would need to create individual glyphs from smaller ones through the use of ligatures, but a similar approach couldn't be taken to Japanese, since many glyphs have multiple meanings depending on context. How should these be represented in Unicode? Yes, these are likely solved problems, but I'm sure there are other examples of less-prominent languages that have similar problems but nobody's put in the work to solve them because the languages aren't as popular online.

You need to be able to represent the language at all stages of authorship - i.e. Unicode needs to be able to represent a half-written Japanese word somehow (yes, Japanese is a bad example because it has a phonetic alphabet as well as a pictograph alphabet).

Anyway, trying to figure out a single text encoding scheme capable of representing every language on Earth is not an easy task.

Post reply on HN