Live data from Hacker News

I Can’t Write My Name in Unicode

modelviewculture.com

281–290 of 377 posts

Re: I Can’t Write My Name in Unicode

#281

Earlier quoted context omitted.

> You don't want a in German and a in English to be different letters just because Helvetica and Baskerville look different. A much more apt comparison would be Antiqua and Fraktur. It was a commonly-held belief that Fraktur was the authentic German alphabet, separate from the alphabets used by other languages. Back in the 19th century, English-German dictionaries even used Antiqua for English words and Fraktur for G…

no, it isn't a more apt comparison. Antiqua and Fraktur look much more different from each other than most Chinese and Japanese characters are from eachother. Here are two screenshots in Japanese and Chinese, taken from the http://www.yomiuri.co.jp/ for one, and http://www.cntv.cn/ for the other. Neither media outlet can be accused of being cultural sellouts misrepresenting the written tradition of their countries. J…

Your screenshots are showing the exact same glyphs because you don't have the "宋体" font installed that's specified by the Chinese web page (it's a script font so you'd notice instantly it's different), so your browser picked a Japanese font to render it instead.

The only reason they look the same in your screenshots is due to Han Unification. Your reasoning is "they're unified in Unicode which is proof they should be unified in Unicode".

Re: I Can’t Write My Name in Unicode

#282
post #271

Earlier quoted context omitted.

.Net seems to do the same thing, Javascript (according to jsfiddle) as well. So maybe this is more widespread than I thought (again - I have never seen that character in the wild)? Java (as in Try Clojure) seems to do the 'expected' SS thing. Trying the golang playground I get even worse: fmt.Println(strings.ToUpper("ßẞ")) returns ßẞ (yeah, unchanged?) So, while I agree that you're technically correct (ẞ exists!) I d…

I think this is more related to the fact that there aren't many sane libraries implementing unicode and locales -- so you'll get either some c lib/c++ lib, system lib, java lib -- or an actual new implementation that's actually been done "seriously" -- as part of being able to say: "Yes, X does actually support unicode strings.". Python3 got a lot of flac for the decision to break away from it's byte sequences, to it…

Thanks for pointing that out -- I was vaguely aware 3.2 wasn't good (but pypy still isn't up to 3.4?) -- it's what's (still) in Debian stable as python3 though. Jessie (soonish to be released) will have 3.4 though, so at that point python3 should really start to be viable (to the extent that there are differences that actually are important...).

For the record, .casefold():

    #Python 3.4:
    >>> 'Åßẞ'.casefold() == 'åßß'.casefold() == 'åssss'
    True
[ed: Also, wrt upper/lower being for display purposes -- I thought it was nice to point out that they are not symmetric, as one might expect them to (although that expectation is probably wrong in the first place...]

Re: I Can’t Write My Name in Unicode

#283

Earlier quoted context omitted.

Ä and Æ are more different from eachother than the characters in Chinese and Japanese which have been merged. They are used for the same thing, they share etymology, but they are not the same letter. By the way, Unicode is about scripts, not languages. If we started distinguishing by language, we might need to start remembering that china doesn't have a single language. Duplicating all those characters again to cover…

> Ä and Æ are more different from eachother than the characters in Chinese and Japanese which have been merged. > They are used for the same thing, they share etymology, but they are not the same letter. That second quote applies equally to Ä and Æ.

It applies to Ä and Æ... which is what the parent said. It doesn't apply to 気 and 气 and 氣. Those are all the same thing.

Re: I Can’t Write My Name in Unicode

#284
post #227
post #156

It sounds like the author is looking to be offended. Talking about Bengali being the seventh largest native language, then saying that US$18k is too expensive for a stake in solving this problem? Emojis with skin tones, something every human can use, that have arrived before a character that only Bengalis use is taken as an " outright insult "?

> Talking about Bengali being the seventh largest native language, then saying that US$18k is too expensive for a stake in solving this problem? West Bengal and Bangladesh aren't exactly the richest places in the world. > Emojis with skin tones, something every human can use, that have arrived before a character that only Bengalis use is taken as an "outright insult"? Nobody is being inconvenienced by the inability t…

I work with a Bangladeshi, and he says that there are more millionaires in Bangladesh than here in Australia. Yes, the people are poor on average, but there's still a lot of money sloshing around.

As to your second point, the top-voted comment in the thread talks about writing exactly that character and gives an example. From the ensuing discussion with the article author, it seems that rather than an inconvenience per se, it's more of a subtle issue. sdg1 also mentions that there's little in the way of local standards, which may be hindering a formal uptake by unicode.

Ultimately the answer is to push the issue with the standards body. If you aren't going to front the money for voting power, then pester the people who do have the power, and make it easier for them by helping them to standardise the codepoints for the language. But whichever path is taken, they must be engaged professionally - saying things like "I take it as an outright insult that you haven't accommodated me yet" is not going to help.

Re: I Can’t Write My Name in Unicode

#285
post #188

Earlier quoted context omitted.

> > It seems like you're trying to single out combining pairs as "less legitimate" when they're extensively used in the standard. > I'm saying that Unicode only does it in English where it makes semantic sense to a native English speaker. Well, combining characters almost never come up in English. The best I can think of would be the use of cedillas, diaereses, and acute accents in words like façade, coördinate and…

> ch, ll, ñ, and rr are all considered separate letters In Spanish "rr" has never been considered as a single letter. "Ch" and "ll" used to be, but not anymore. Ñ is, of course.

I think i'm older than you. I learnt in the school they were different letters, and also I remember when they were removed at the beginning of 90's

Re: I Can’t Write My Name in Unicode

#286
post #216

Earlier quoted context omitted.

>" Was British rule actually a net negative, in retrospect? Have there been studies done using objective criteria (not emotional) over counties that were colonies versus ones that weren't? " It's both a positive, and a negative. But if you want to calculate whether it was a "net" positive or "net" negative, then I'm afraid you're going to have to quantify very difficult-to-quantify concepts. And you'll have to do it…

Well that's why I said without taking into account people that were destroyed by it. That is, if we look purely at modern-day circumstances, are they better? A comparison might be a country with overpopulation. If you kill x% of the population and enact some birth control plan, then 100 years later you might end up with a "net positive" discounting the people that were killed and the emotional effects of their loved…

http://www.quora.com/Which-countries-in-Africa-were-never-co...

Alternatively, you could compare countries that had "more" or "less" colonialist control/influence, and see if there is any correlation to their current performance. Though, as with all social-sciency it's very difficult to isolate the variables to get deep insight into possible causal links.

Re: I Can’t Write My Name in Unicode

#287
post #117

Earlier quoted context omitted.

> The Bengali fonts that you have installed happen to render it as a jophola It's not only the Bengali font - the text rendering framework of my operating system also needs to have a bunch of complex rules to figure out that a jophola needs to be rendered. It also needs to know that the visual ordering of i-kar is before the preceding consonant cluster (দ in আদিত্য). > the characters that are required to type a jopho…

Is Bengali your first language? While one can make the case that ত্য is simply "'to' - 'o' + 'ya' = 'to'"[0][1], it's rather confusing mental acrobatics, and it doesn't reflect either how the writing system is taught, or how native speakers use it and think of it on a day-to-day basis. If anything, your comment makes a stronger argument for consolidating ই and ি (they are literally the same letter and phoneme, but wr…

It seems to me that the high-level issue here is that Unicode is caught between people who want it to be a set of alphabets, and people who want it to be a set of graphemes.

The former group would give each "semantic character" its own codepoint, even when that character is "mappable" to a character in another language that has the same "purpose" and is always represented with the same grapheme (see, for example, latin "a" vs. japanese full-width "a", or duplicate ideograph sets between the CJK languages.) In extremis, each language would be its own "namespace", and a codepoint would effectively be described canonically as a {language, offset} pair.

The latter group, meanwhile, would just have Unicode as a bag of graphemes, consolidated so that there's only one "a" that all languages that want an "a" share, and where complex "characters" (ideographs, for example, but what we're talking about here is another) are composed as ligatures from atomic "radical" graphemes.

I'm not sure that either group is right, but trying to do both at once, as Unicode is doing, is definitely wrong. Pick whichever, but you have to pick.

Re: I Can’t Write My Name in Unicode

#288
post #262

Earlier quoted context omitted.

Not necessarily disagreeing with your broader point but I just want to point out that the examples are only obscure and antiquared in English. æ is common in modern Danish and unambiguously a separate letter, as is þ in Icelandic.

And Norwegian (for æ). We have æ/Æ and ø/Ø (distinct, in theory, from the symbol for the empty set, btw), while å/Å used to be written aa/AA a long time ago (but is obviously not a result of combining two a's). Swedish uses ö/Ö for essentially ø/Ø, ä/Ä for æ/Æ. And both of those I can only easily type by combining the dot-dot, with o/O, a/A because my Norwegian keyboard layout has key for/labelled øæå, not öäå. For J…

To be fair, proper codepoint processing is a pain even in Java, which was created back when Unicode was in 16-bit mode. Now that it's extended to 32-bits, proper Unicode string looping looks something like this:

    for(int i = 0; i 

Re: I Can’t Write My Name in Unicode

#289
post #156

It sounds like the author is looking to be offended. Talking about Bengali being the seventh largest native language, then saying that US$18k is too expensive for a stake in solving this problem? Emojis with skin tones, something every human can use, that have arrived before a character that only Bengalis use is taken as an " outright insult "?

>Emojis with skin tones, something every human can use, that have arrived before a character that only Bengalis use is taken as an "outright insult"?

Who cares if emojis can be used by every human? Bengali language support is something people actually need to to function normally in the digital age without jumping through a million hoops.

Emojis with skin tones are a joke.

Re: I Can’t Write My Name in Unicode

#290

Earlier quoted context omitted.

Han unification as a hole is misguided? I'll grant you that some characters which were unified probably shouldn't have been, and maybe some that some that should have been weren't, but what's the argument for the whole thing to be misguided? Should Norwegian A and English A be different Unicode code points just because Norwegian also has Ø, proving that it is a different writing system? You may want to debate whether…

We'll the Turkish i/ı/I/I is I think exactly the example I would have come up with of characters that looks the same as i/I, but should have it's own code point, just like cyrillic characters have their own code points despite looking like latin characters.

Absolutely. So i/ı/I/I do have their own codepoints. But the rest of the letters, which are the same, don't. Just like han unification. Letters which are the same are the same, and those which are not are not, even if they look pretty close.
Post reply on HN