Live data from Hacker News

I Can’t Write My Name in Unicode

modelviewculture.com

141–150 of 377 posts

Re: I Can’t Write My Name in Unicode

#141
I think the author misses the point completely.

Things like this: "No native English speaker would ever think to try “Greco Unification” and consolidate the English, Russian, German, Swedish, Greek, and other European languages’ alphabets into a single alphabet."

The author probably ignores that different European languages used different alphabet scripts until very recently. For example, Gothic and other different script were used in lots of books.

I have old books that take something like a week to be able to read fast, and they are in German!.

But it was chaos and it unified into a single script. Today you open any book, Greek, Russian, English or German and they all use the same standard script, although they include different glyphs. There is a convention for every symbol. In cyrilic you see an "A" and a "a".

In fact, any scientific or technical book includes Greek letters as something normal.

It should also be pointed out that latin characters are not latin, but modified latin. E.g Lower case letters did not exist on Roman's empire. It were included by other languages and "unified".

Abut CJK, I am not an expert but I had lived on China, Japan and Korea, and in my opinion it has been pushed by the Governments of those countries because it has lots of practical value for those countries.

Learning Chinese characters is difficult enough. If they don't simplify it people just won't use it, when they can use Hangul or kana. With smartphones and tablets people there are not hand writing Chinese anymore.

It makes no sense to name yourself with characters nobody can remember.

Re: I Can’t Write My Name in Unicode

#143
post #104
post #73

Earlier quoted context omitted.

Imagine a world where the British always write the lowercase letter g as a single-story glyph ( http://en.m.wikipedia.org/wiki/G#Typographic_variants ). The colonies start writing it identically, but after a while, they start writing it as a double-story g. After a century or so, nobody in he colonies writes the single-story variant, and all Brits always do. The unicode consortium studies the case and concludes that…

No, you don't. It's the same letter and it's always "goto". Your local setup determines the look of the glyph, so nobody sees an unfamiliar form. But maybe you'd like to encode typefaces/fonts in the Unicode code points, as well? To make sure that I'm seeing the exact same arrangement of pixels you want me to see?

This is the core of the Han Unification debate.

"G" and "g" are the same letter. They started off as stylistic forms of a unicameral alphabet. Over time they took on separate meanings, and now we have a bicameral alphabet, where the two forms have different code points.

Of course, over the last 2000 years, we've developed rules for how to use them. "I was reading a nice book on Polish polish on the way from Reading to Nice" contains three pairs of words where the capitalization changes the meaning and pronunciation. (In simplified form, "What do you know about polish?" is different than "What do you know about Polish?")

If there were a simple rule to specify capitalization, eg, only the first letter of a sentence, and it were easy to detect the start of a sentence, then the alternate you might say it's pointless to have both "g" and "G"; we should have only a single form and let the local setup determine how to display it. (Something like the Greek sigma, which has the form ς when used at the end of a word, though Unicode has them as two different glyphs.)

In Someone's nice example, it's easy to think of how the two divergent forms of 'g' might take on their own meaning. Perhaps the Americans have decided that double-story g was the sign of true patriots, and that single-story g was for traitors. (Akin to the shibboleth of how to say 'H' in Northern Ireland; aitch was Protestant, haitch was Catholic, and using the wrong version could get you into trouble.) Perhaps they started to use the new 'g' preference as a currency symbol, in the way that £ is the same letter as L, from the Latin libra pondo.

Re: I Can’t Write My Name in Unicode

#144

I am an Indian and it shocks me that Indians are still blaming the British after 70 yrs of independence. Is 70 years of Independence not enough to make your language "first class citizen" ? Ofcourse Bengali is second class language because Bengalis didn't invent the standard. Can we stop blaming white people for everything. Seriously WTF.

Was British rule actually a net negative, in retrospect? Have there been studies done using objective criteria (not emotional) over counties that were colonies versus ones that weren't? I suppose you can't really quantify the value of people that were destroyed by colonization, but you can look at the current population.

Also I just gotta wonder: suppose European or other relatively simple-to-encode languages didn't exist, and everyone used the OPs language. How would they have handled advancing computers? I've seen photos of Japanese typewriters and they look... unwieldy to say the least. And graphics tech took a while to get advanced enough to handle such languages, let alone input. (MS Japanese IME uses some light AI to pick the desired kanji, right?)

Disclaimer: I don't mean this in an offensive way, just a dispassionate curiosity.

Re: I Can’t Write My Name in Unicode

#145

I think the author misses the point completely. Things like this: "No native English speaker would ever think to try “Greco Unification” and consolidate the English, Russian, German, Swedish, Greek, and other European languages’ alphabets into a single alphabet." The author probably ignores that different European languages used different alphabet scripts until very recently. For example, Gothic and other different s…

Right, I was going to argue against that too. Changing fonts every letter will always look weird. There's a lot of different ways to shape these letters that are all counted as the same.

Re: I Can’t Write My Name in Unicode

#146
post #70

I came in expecting to read an article bemoaning some niche language and playing the diversity card. I was not disappointed, but as I kept reading, the author made some very good points. I don't really care that the organization is run by white men who speak English, because frankly the entire computing industry and telecommunications industry is based on that. I'm not going to argue about the original sin there, bec…

Han unification has been overly aggressive about merging some characters, but the basic principle is not as flawed as it is some times (as in this article) made to sound. The vast majority of Japanese and Chinese characters are not only similar, they are identical. Not all are. Some are clearly different characters deriving from a common historical root, and should not be unified. Sometimes, when characters are a bit…

> You don't want a in German and a in English to be different letters just because Helvetica and Baskerville look different.

A much more apt comparison would be Antiqua and Fraktur. It was a commonly-held belief that Fraktur was the authentic German alphabet, separate from the alphabets used by other languages. Back in the 19th century, English-German dictionaries even used Antiqua for English words and Fraktur for German words.

Re: I Can’t Write My Name in Unicode

#147
post #143
post #104

Earlier quoted context omitted.

No, you don't. It's the same letter and it's always "goto". Your local setup determines the look of the glyph, so nobody sees an unfamiliar form. But maybe you'd like to encode typefaces/fonts in the Unicode code points, as well? To make sure that I'm seeing the exact same arrangement of pixels you want me to see?

This is the core of the Han Unification debate. "G" and "g" are the same letter. They started off as stylistic forms of a unicameral alphabet. Over time they took on separate meanings, and now we have a bicameral alphabet, where the two forms have different code points. Of course, over the last 2000 years, we've developed rules for how to use them. "I was reading a nice book on Polish polish on the way from Reading t…

Absolutely. But the two "g" haven't diverged, yet.

We don't give out code points to speculative future developments.

If and when they diverge one will get its very own code point.

Re: I Can’t Write My Name in Unicode

#148
post #106

Earlier quoted context omitted.

Unicode makes extensive use of combining characters for european languages, for example to produce diacritics: ìǒ or even for flag emoji. A correct rendering system will properly combine those, and if it doesn't then that's a flaw in the implementation, not the standard. It seems like you're trying to single out combining pairs as "less legitimate" when they're extensively used in the standard.

> Unicode makes extensive use of combining characters for european languages, for example to produce diacritics: ìǒ or even for flag emoji. But it doesn't, for example say that a lowercase "b" is simply "a lowercase 'l' followed by an 'o' followed by an invisible joiner", because no native English speaker thinks of the character "b" as even remotely related to "lo" when reading and writing. > It seems like you're try…

> > It seems like you're trying to single out combining pairs as "less legitimate" when they're extensively used in the standard.

> I'm saying that Unicode only does it in English where it makes semantic sense to a native English speaker.

Well, combining characters almost never come up in English. The best I can think of would be the use of cedillas, diaereses, and acute accents in words like façade, coördinate and renownèd (I've been reading Tolkien's translation of Beowulf, and he used renownèd a lot).

Thinking about the Spanish I learned in high school, ch, ll, ñ, and rr are all considered separate letters (i.e., the Spanish alphabet has 30 letters; ch is between c and d, ll is between l and m, ñ is between n and o, and rr is between r and s; interestingly, accented vowels aren't separate letters). Unicode does not provide code points for ch, ll, or rr; and ñ has a code point more from historical accident than anything (the decision to start with Latin1). Then again, I don't think Spanish keyboards have separate keys for ch, ll, or rr.

Portuguese, on the other hand, doesn't officially include k or y in the alphabet. But it uses far more accents than Spanish. So, a, ã and á are all the same letter. In a perfect world, how would Unicode handle this? Either they accept the Spanish view of the world, or the Portuguese view. Or, perhaps, they make a big deal about not worrying about languages and instead worrying about alphabets ( http://www.unicode.org/faq/basic_q.html#4 ).

They haven't been perfect. And they've certainly changed their approach over time. And I suspect they're including emoji to appear more welcoming to Japanese teenagers than they were in the past. But (1) combining characters aren't second-class citizens, and (2) the standard is still open to revisions ( http://www.unicode.org/alloc/Pipeline.html ).

Re: I Can’t Write My Name in Unicode

#149
post #101

Earlier quoted context omitted.

Antiqua, Fractura, Schwabacher, Textura and all those other charming variants of writing European languages don't have seperate code points for all those presentational variations. Do we white Western men happen to discriminate against ourselves?

If you don't know the specifics, why the hell are you even arguing this point?

I don't know what?

If you're just here to randomly insult people, please leave.

Update: ah, you just mishandled the threading.

But still... I asked very specific questions, in order to understand your point, and you're very combative. My point stands.

Re: I Can’t Write My Name in Unicode

#150
Out of curiosity, before Unicode came along what was the state of the art for encoding/writing/displaying Bengali?

Sometimes I think issues with Unicode might be because it's trying to solve issues for languages that haven't yet arrived at a good solution for themselves yet.

Latin-using languages ended up going through a very long orthographic simplification and unification process after centuries of divergent orthographic development. These changes all occurred to simplify or improve some aspect of the language: improve readability, increase typesetting speed, reduce typesetting pieces. Early personal computers even did away with lowercase letters and some punctuation marks completely before they were grudgingly reintroduced.

I'm more familiar with Hangul (Korean), which has sometimes complex rules for composing syllables but has undergone fairly rapid orthographic changes once it was accepted into widespread and official use: dropping antiquated or dialectal letters, dropping antiquated tone markers, spelling revisions, etc. In addition, Chinese characters are now completely phased out in North Korea and are rapidly disappearing in the South.

It's my personal theory that the acceleration of this orthographic change has had to do with the increased need to mechanize (and later computerize) writing. Koreans independently came up with workable solutions to encode, write and display their system, sometimes by changing the system, sometimes by figuring out how to mechanically and then algorithmically encode the complex rules for their script. It appears that a happy medium has been found and Korean users happily pound away at keyboards for billions of keystrokes every year.

I'm digressing here, but pre-unicode, how had Bengali users solved the issue of mechanically and algorithmically encoding, inputting and displaying their own language? Is it just a case that the unicode team is ignorant of these solutions and hasn't transferred this hard-earned knowledge over?

(note, I came across this page as I was writing this, how suitable was this solution? http://www.dsource.in/resource/devanagari-letterforms-histor...)

I'm asking these questions out of ignorance of course, I don't know the story on either side.

On a second point, I'm deeply concerned about Unicode getting turned into a font-repository for slightly different (or even the same) character, that just happens to end up in another language. For example, Cherokee uses quite a few Latin letters (and numbers) (and quite a few original inventions). Is it really necessary to store all the Latin letters it uses again? Would a reader of Tsalagi really care too much if I used W or Ꮃ? When does a character go from being a letter to being a specific depiction of that letter?

Post reply on HN