Live data from Hacker News

I Can’t Write My Name in Unicode

modelviewculture.com

91–100 of 377 posts

Re: I Can’t Write My Name in Unicode

#91
post #77
post #72

Earlier quoted context omitted.

I don't understand Bengali at all. I'm trying to understand your second sentence though. When you say "no one writes that way", do you mean nobody hits the keys for letter, followed by vowel-silencing diacritic, followed by another vowel? Or do you mean the glyph that results from that combination of keystrokes doesn't match how a Bengali speaker would write it on paper? If it's the latter, isn't that an issue for th…

It has to do with how the text is rendered. For example, if you see the Bengali text on page 3 of this PDF: http://www.unicode.org/L2/L2004/04252-khanda-ta-review.pdf it is unreadable and incorrect Bengali. ;-)

But Unicode emphatically does not define a rendering, a glyph.

To me it sounds like the rendering needs to be fixed, not Unicode.

Re: I Can’t Write My Name in Unicode

#92
post #48

Not sure if the l33tspeak analogy is fully justified. In case of the "missing" letter (called khanda-ta in Bengali) for the Bengali equivalent of "suddenly", historically, it has been a derivative of the ta-halant form (ত + ্ + ‍ ). As the language evolved, khanda-ta became a grapheme of its own, and Unicode 4.1 did encode it as a distinct grapheme. A nicely written review of the discussions around the addition can b…

> I could write the author's name fine: আদিত্য

Author here.

Well, yes and no. The jophola at the end is not actually given its own codepoint[0]. The best analogy I can give is to a ligature in English[1]. The Bengali fonts that you have installed happen to render it as a jophola, the way some fonts happen to render "ff" as "ff" but that's not the same thing as saying that it actually is a jophola (according to the Unicode standard).

The difference between the jophola and an English ligature, though, is English ligatures are purely aesthetic. Typing two "f" characters in a row has the same obvious semantic meaning as the ff ligature, whereas the characters that are required to type a jophola have no obvious semantic, phonetic, or orthographic connection to the jophola.

[0] http://unicode.org/charts/PDF/U0980.pdf

[1] Some fonts will render (e.g.) two "f"s in a row as if they were a ligature, even though it's not a true ff(U+FB00).

Re: I Can’t Write My Name in Unicode

#93
post #54

“Whatever path we take, it’s imperative that the writing system of the 21st century be driven by the needs of the people using it. In the end, a non-native speaker – even one who is fluent in the language – cannot truly speak on behalf the monolingual, native speaker.” Not sure how the author can simultaneously say this, while criticizing the CJK unification, which makes total sense, and has never been a point of con…

> [...] CJK unification[...] has never been a point of contention in the communities concerned with it. I am not very familar with the CJK unification project so take my points with a grain of salt. > more than not opposing CJK unification, I benefit from it greatly. I think think that is a different point of view. Isn't it? You are seeing your benefit whereas the author is seeing his. Here's an alternative solution:…

No experience in Asian languages, but J is pronounced differently in English and German. Should they have unique characters?

Re: I Can’t Write My Name in Unicode

#94
> He proudly announces that there are ‘no fewer than 147 Indian dialects’ – a pathetically inaccurate count.

Wow. How can a country function like this? Is everyone proficient in their native language plus a 'common' one, or are all interactions supposed to be translated inside the same country? Regardless of historical and cultural value, if that's the case, it seems... inefficient.

I do realize that there are more countries like this, but the number of languages seems way too high. I am really curious how that works.

Re: I Can’t Write My Name in Unicode

#95
post #9

Earlier quoted context omitted.

So blame the indian government for any problems with ISCII. The problems with Unicode support for global languages are indeed to be blamed on the Unicord Consortium. Nothing the indian governemnt did prevents the UC from availing themselves of the global knowledge needed to create a good global standard.

do you have any evidence that the UC has actively ignored requests from Bengali speakers? Has any Bengali speaker made proposals to the UC for fixing these issues? If yes, and the UC chose to ignore them, then there is some blame to be assigned with the UC. Otherwise, this is a non-issue. Take, for example, Tibetan. The number of Tibetan speakers is minuscule compared to, say, Bengali. But still Tibetan has good supp…

I think it needs to be stressed that anyone can submit character proposals to the consortium and work with them to get them included in the next versions of the standard. You don't need to be a member (By paying the hefty fee. Someone needs to pay for the operational costs of the consortium.) to have your suggestions taken into account.

Re: I Can’t Write My Name in Unicode

#96
post #64

I am an Indian and it shocks me that Indians are still blaming the British after 70 yrs of independence. Is 70 years of Independence not enough to make your language "first class citizen" ? Ofcourse Bengali is second class language because Bengalis didn't invent the standard. Can we stop blaming white people for everything. Seriously WTF.

Hebrew (and I'd guess Arabic and other right-to-left languages) work rather badly in Unicode when it comes to bidirectional rendering; however, to the extent that it's a result of Israeli/Egyptian/Saudi/etc. companies and/or governments failing to pay $18K (the figure from TFA) to pay for the consortium membership, I kinda blame them and not the consortium. I mean, it's not a whole lot of money, even for a smallish c…

The pile of poo is a great example. It is supported because the unicode consortium pays attention to the need of non western languages, not because it ignores them and would rather do silly things.

Japan has had a tendency to includes all sorts of crap, including this example of actual crap, into its character sets/encodings. Japan also has a history of working with the various standard bodies like the Unicode Consortium. Because Japanese systems had the pile of poo, unicode has it.

Re: I Can’t Write My Name in Unicode

#97
post #55

Wait, "ত + ্ + ‍ = ‍ৎ" is nothing like "\ + / + \ + / = W". The Bengali script is (mostly) an abugida. Ie, consonants have an inherent vowel (/ɔ/ in the case of Bengali), which can be overriden with a diacritic representing a different vowel. To write /t/ in Bengali, you combine the character for /tɔ/, "ত", with the "vowel silencing diacritic" to remove the /ɔ/, " ্". As it happens, for "ত", the addition of the diacr…

I don't understand Bengali at all, but the character ৎ does have it's own Unicode codepoint (U+09CE, BENGALI LETTER KHANDA TA). It was introduced in Unicode 4.1 in 2005.

Yes, this is what the article says:

> Until 2005, Unicode did not have one of the characters in the Bengali word for “suddenly”.

The codepoint does exist now, but it took ten years before it was included.

Re: I Can’t Write My Name in Unicode

#98

Am I the only person who thought unifying the Greco-Roman language characters actually sounds like a good idea?

Yeah, I thought that "No native English speaker would ever think to try “Greco Unification”" was a poor argument. In seems like a reasonable idea.

I think their argument was: The characters look the same (e.g. Russian's first character and the English A) but have different meanings.

So in this example if you searched for the English word "Eat" that is also a completely legal Russian word (E, A, and T, exist in English and Russian), however it means nothing remotely similar.

I don't know if they're right or wrong. I am just saying that might be the point they were trying to make. You could make a Greco Unified unicode set and it would work fairly well, but you might wind up with some confusing edge cases where it isn't clear what language you're reading (literally).

This could be particularly problematic for automation (e.g. language detection). Since in some situations any Greco-like language could look similar to any other (in particular as the text gets shorter).

Re: I Can’t Write My Name in Unicode

#99

“Whatever path we take, it’s imperative that the writing system of the 21st century be driven by the needs of the people using it. In the end, a non-native speaker – even one who is fluent in the language – cannot truly speak on behalf the monolingual, native speaker.” Not sure how the author can simultaneously say this, while criticizing the CJK unification, which makes total sense, and has never been a point of con…

I agree, but it seems tricky like it would be tricky to strike exactly the right balance between unifying too much and too little.

The arguments for Han unification could have just as well been applied to unifying the Nordic languages – the Swedish Ä is really exactly the same letter as Danish/Norwegian Æ (except that Swedish words never ever use the latter and presumably Danish words never use the former), so it could be argued that it should have the same codepoint, forcing us to use country-specific fonts and making it impossible to use both languages in a simple text file.

So, for example, running "git log" on a project with an international contributor base would render either Swedish och Danish names incorrectly, depending on which font the user has selected. That would look extremely strange to me, so I imagine Chinese/Japanese/Korean users feel similarly.

(Luckily for us Nordic users, our letters were already separate characters in ISO-8859-1, which Unicode is a superset of.)

Re: I Can’t Write My Name in Unicode

#100
post #73

Earlier quoted context omitted.

Isn't the idea technically that the code shouldn't even have to guess? Why isn't this the case?

Imagine a world where the British always write the lowercase letter g as a single-story glyph ( http://en.m.wikipedia.org/wiki/G#Typographic_variants ). The colonies start writing it identically, but after a while, they start writing it as a double-story g. After a century or so, nobody in he colonies writes the single-story variant, and all Brits always do. The unicode consortium studies the case and concludes that…

Isn't this the question of a font? In which case the client chooses if they want to use a font with a double-story or a single-story g?
Post reply on HN