Live data from Hacker News

I Can’t Write My Name in Unicode

modelviewculture.com

121–130 of 377 posts

Re: I Can’t Write My Name in Unicode

#121
post #16
post #4

I wonder if the author has submitted a proposal to get the missing glyph for their name added. You don't need to be a member of the consortium to propose adding a missing glyph/updating the standard. The point of the committee as I understand it isn't to be an expert in all forms of writing, but to take the recommendations from scholars/experts and get a working implementation, though more diverse representation of l…

Holy cow CJK unification is a terrible idea. Maybe if it originated from the CJK governments, it might be an OK idea, but the idea of a Western multinationals trying to save Unicode space by disregarding the distinctness of a whole language group is idiotic. The fundamental roll of an institution like the Unicode Consortium is to be descriptive, not prescriptive. If there is a human script, passing certain, low, low…

The Han unification working group (IRG) members were in fact from Asia and appointed by their governments. Why would you think otherwise?

Apparently this even includes North Korea.

Re: I Can’t Write My Name in Unicode

#122
post #101
post #87

Earlier quoted context omitted.

> The vast majority of Japanese and Chinese characters are not only similar, they are identical. Not all are. Some are clearly different characters deriving from a common historical root, and should not be unified. And what about traditional versus simplified? Which glyph set do I use? Oh wait, thanks to Han unification, I now need to rely on bloody environment variables to decide! For Chinese text, rendering a strin…

Antiqua, Fractura, Schwabacher, Textura and all those other charming variants of writing European languages don't have seperate code points for all those presentational variations. Do we white Western men happen to discriminate against ourselves?

Those are all stylistic differences. You can still read any of those fonts. Sure some designer can go overboard and make a font hard to read, but that is just a designer being overly fancy!

Think instead of the difference between Cyrillic and Latin.

Sure if you squint hard enough they both have a common origin (Greek), but you'd be rather annoyed if you setup your phone, selected "English" and the OS used Cyrillic letters to spell out English words.

Likewise you'd be annoyed if Greek letters were used.

And you'd be even more upset if some products decided to use Cyrillic, some Greek, and some used Latin.

Re: I Can’t Write My Name in Unicode

#123

Earlier quoted context omitted.

Yeah, I thought that "No native English speaker would ever think to try “Greco Unification”" was a poor argument. In seems like a reasonable idea.

I think their argument was: The characters look the same (e.g. Russian's first character and the English A) but have different meanings. So in this example if you searched for the English word "Eat" that is also a completely legal Russian word (E, A, and T, exist in English and Russian), however it means nothing remotely similar. I don't know if they're right or wrong. I am just saying that might be the point they we…

English, French, German, Italian, Spanish and several other European languages have mostly identical character sets and even large numbers of similar or identical words. Computers detect these languages just fine. I think we'll be okay.

Re: I Can’t Write My Name in Unicode

#124
post #5

I get the author's point, but assigning blame to the Unicode Consortium is incorrect. The Indian government dropped the ball here. They went their separate way with "ISCII" and other such misguided efforts, instead of cooperating with UC. To me, the UC is just a platform. The government is the de facto safeguarder of the people's interests; if it drops the ball, it should be taken to task, not the provider of the pla…

The world (especially the technological) is changing too fast to rely on nation-state channels that proceed by forming commissions and writing position statements blah blah blah. It would be better for the Unicode consortium to start actively soliciting input and/or technical contributions from people.

Re: I Can’t Write My Name in Unicode

#125
post #17

> He proudly announces that there are ‘no fewer than 147 Indian dialects’ – a pathetically inaccurate count. (Today, India has 57 non-endangered and 172 endangered languages, each with multiple dialects – not even counting the many more that have died out in the century since My Fair Lady took place) So, how many were there really? At the time, I mean.

I think the numbers are somewhat disputed. The People's Linguistic Survey of India says there are at least 780, with ~220 having died out in the last half century.[1] The Anthropological Survey of India reported 325 languages.[2] The discrepancies are made particularly tricky because of the somewhat ambiguous distinction between languages and dialects. [1] http://blogs.reuters.com/india/2013/09/07/india-speaks-780-l.…

So, this person thinks there are 229 languages in India, but actually there's a minimum of 325, possibly twice that?

Re: I Can’t Write My Name in Unicode

#126
post #104
post #73

Earlier quoted context omitted.

Imagine a world where the British always write the lowercase letter g as a single-story glyph ( http://en.m.wikipedia.org/wiki/G#Typographic_variants ). The colonies start writing it identically, but after a while, they start writing it as a double-story g. After a century or so, nobody in he colonies writes the single-story variant, and all Brits always do. The unicode consortium studies the case and concludes that…

No, you don't. It's the same letter and it's always "goto". Your local setup determines the look of the glyph, so nobody sees an unfamiliar form. But maybe you'd like to encode typefaces/fonts in the Unicode code points, as well? To make sure that I'm seeing the exact same arrangement of pixels you want me to see?

The issue is that, in this example, both the colonials and the British think the two 'g' characters are different letters, just as people nowadays think 'g' and 'G' are different characters (historically, that is at least up for debate: http://en.m.wikipedia.org/wiki/Letter_case#History). Americans will want to see a different character when quoting Shakespeare inside American-English text (globalization starts with a different letter than globalisation)

because Unicode has only one character, writing that text in a text editor or storing it in a text column in a database becomes impossible.

Yes, there are workarounds such as using escape characters or html, but those are a nuisance that could be avoided by including both variants in Unicode.

The unicode consortium is entitled to think differently, but they cannot expect everybody to be happy with their choice.

Re: I Can’t Write My Name in Unicode

#127
post #67

Earlier quoted context omitted.

Nor is it unreasonable to "unify" Latin, Greek and Cyrilic: Cyrillic ПФ vs Greek ΠΦ Cyrillic АВ vs Latin AB Obviously using ω for w (as he does) is stupid, but his reducto-ad-absurdum is not particularly absurd.

Unicode kind of does this already with dotless 'i'; capital 'ı' and lowercase 'İ' are represented as regular latin 'I' and 'i' respectively, despite being semantically different letters.

hungarian "a" is also a separate letter from hungarian "á" but shares the same glyph with english "a" (edit: all vowels are actually considered different letters in their accented form in hungarian, while obviously they are the same letter with a modifier in some latin languages).

Re: I Can’t Write My Name in Unicode

#128
post #101
post #87

Earlier quoted context omitted.

> The vast majority of Japanese and Chinese characters are not only similar, they are identical. Not all are. Some are clearly different characters deriving from a common historical root, and should not be unified. And what about traditional versus simplified? Which glyph set do I use? Oh wait, thanks to Han unification, I now need to rely on bloody environment variables to decide! For Chinese text, rendering a strin…

Antiqua, Fractura, Schwabacher, Textura and all those other charming variants of writing European languages don't have seperate code points for all those presentational variations. Do we white Western men happen to discriminate against ourselves?

The difference between traditional and simplified Chinese characters is more than simply different fonts. Part of the difficulty is that some simplified characters map to multiple traditional characters, which means that converting from one to the other may be lossy. There's also the Japanese equivalent of simplified characters (shinjitai), many of which differ from their Chinese counterparts, as well as characters that were invented in Japan and may or may not have Chinese equivalents (kokuji).

Re: I Can’t Write My Name in Unicode

#129
post #101

Earlier quoted context omitted.

Antiqua, Fractura, Schwabacher, Textura and all those other charming variants of writing European languages don't have seperate code points for all those presentational variations. Do we white Western men happen to discriminate against ourselves?

Those are all stylistic differences. You can still read any of those fonts. Sure some designer can go overboard and make a font hard to read, but that is just a designer being overly fancy! Think instead of the difference between Cyrillic and Latin. Sure if you squint hard enough they both have a common origin (Greek), but you'd be rather annoyed if you setup your phone, selected "English" and the OS used Cyrillic le…

Actually, I can't. I can decipher quite a bit with lots of effort, but I wouldn't call that "reading".

And there are very few people around who can read all those (especially Textura) without problems.

On your other example: when I was in Russia, I found those "unknown" letters difficult, but fun. But the letters they share with Latin? I didn't see a difference.

And to your phone example: of course, I'd be annoyed. But that's exactly the point: you have to set up your local system correctly, to your expectations and standards.

I'd be much more annoyed if some web page showed an English text in some Chinese transliteration, just because the author (writing English!) was Chinese. But that's basically what you propose!

Sorry, I truly think you've run up an argumentative dead end.

Re: I Can’t Write My Name in Unicode

#130
post #73

Earlier quoted context omitted.

Imagine a world where the British always write the lowercase letter g as a single-story glyph ( http://en.m.wikipedia.org/wiki/G#Typographic_variants ). The colonies start writing it identically, but after a while, they start writing it as a double-story g. After a century or so, nobody in he colonies writes the single-story variant, and all Brits always do. The unicode consortium studies the case and concludes that…

Isn't this the question of a font? In which case the client chooses if they want to use a font with a double-story or a single-story g?

When reading 'gas', the client will have to figure out whether that is about a liquid, in which case it has to choose a colonial 'g' (when written with a British 'g' 'gas' always is a liquid). If the meaning is that of a gaseous substance, the client will have to do additional work to determine what kind of 'g' to write.
Post reply on HN