Live data from Hacker News

I Can’t Write My Name in Unicode

modelviewculture.com

131–140 of 377 posts

Re: I Can’t Write My Name in Unicode

#131

Earlier quoted context omitted.

Nor is it unreasonable to "unify" Latin, Greek and Cyrilic: Cyrillic ПФ vs Greek ΠΦ Cyrillic АВ vs Latin AB Obviously using ω for w (as he does) is stupid, but his reducto-ad-absurdum is not particularly absurd.

Not unifying them means that the fonts automatically work when you mix text/names written in these alphabets. It also means that mathematical/physical/chemical stuff (that typically uses Latin and Greek letters together) will just work. There is a similar reasoning behind all the mathematical alphabets in Unicode. Furthermore, Unicode was supposed to handle transcoding from all important preexisting encodings to Unic…

> Not unifying them means that the fonts automatically work when you mix text/names written in these alphabets. It also means that mathematical/physical/chemical stuff (that typically uses Latin and Greek letters together) will just work.

These are already completely separate symbols. Ignoring precomposition, there are at least 4 different lowercase omegas in unicode: APL (⍵ U+2375 "APL FUNCTIONAL SYMBOL OMEGA"), cyrillic (ѡ U+0461 "CYRILLIC SMALL LETTER OMEGA"), greek (ω U+03C9 "GREEK SMALL LETTER OMEGA") and Mathematics (𝜔 U+1D714 "MATHEMATICAL ITALIC SMALL OMEGA").

Re: I Can’t Write My Name in Unicode

#133

Earlier quoted context omitted.

> For me, many of these problems are more of an input issue, than an encoding issue. I think you've hit the nail on the head here. I'm a native English speaker, so I may in fact be making bad assumptions here, but I think the biggest issue here is that people conflate text input systems with text encoding systems. Unicode is all about representing written text in a way that computers can understand. But the way that…

They're not unrelated though. You have to have a way to get from your input format to the finished product in a consistent way, and the glyph set you design has a large bearing on that. You can't solve it completely with AI, because then you just have an AI interpretation of human language, not human language. A language like Korean written in Hangul would need to create individual glyphs from smaller ones through th…

Its not an AI issue, just a small matter of having lots of rules. Moreover this is not just an issue for non-Western languages: the character â (lower case "a" with a circumflex) can be represented either as a single code-point U+00E2 or as an "a" combined with a "^". Furthermore Unicode implementations are required to evaluate these two versions as being equal in string comparisons, so if you search for the combined version in a document, it should find the single code point instances as well.

Re: I Can’t Write My Name in Unicode

#134
post #117

Earlier quoted context omitted.

> I could write the author's name fine: আদিত্য Author here. Well, yes and no. The jophola at the end is not actually given its own codepoint[0]. The best analogy I can give is to a ligature in English[1]. The Bengali fonts that you have installed happen to render it as a jophola, the way some fonts happen to render "ff" as "ff" but that's not the same thing as saying that it actually is a jophola (according to the Uni…

> The Bengali fonts that you have installed happen to render it as a jophola It's not only the Bengali font - the text rendering framework of my operating system also needs to have a bunch of complex rules to figure out that a jophola needs to be rendered. It also needs to know that the visual ordering of i-kar is before the preceding consonant cluster (দ in আদিত্য). > the characters that are required to type a jopho…

Is Bengali your first language?

While one can make the case that ত্য is simply "'to' - 'o' + 'ya' = 'to'"[0][1], it's rather confusing mental acrobatics, and it doesn't reflect either how the writing system is taught, or how native speakers use it and think of it on a day-to-day basis.

If anything, your comment makes a stronger argument for consolidating ই and ি (they are literally the same letter and phoneme, but written differently in different contexts) than for combining the viram and য into the jophola.

[0] To non-Bengali speakers reading this, yes, this is how that construction would work, and yes, I am aware that the arithmetic doesn't appear to add up (which I guess is part of the point).

[1] Also, now that I think about it, the য is a consonant, not a vowel, so using it in place of a vowel is doubly awkward. This is particularly an issue in Bengali, where sounds that might be consonants in English (like "r" and "l") can be either consonants or vowels in Bengali, depending on the word.

Re: I Can’t Write My Name in Unicode

#135
post #101

Earlier quoted context omitted.

Antiqua, Fractura, Schwabacher, Textura and all those other charming variants of writing European languages don't have seperate code points for all those presentational variations. Do we white Western men happen to discriminate against ourselves?

The difference between traditional and simplified Chinese characters is more than simply different fonts. Part of the difficulty is that some simplified characters map to multiple traditional characters, which means that converting from one to the other may be lossy. There's also the Japanese equivalent of simplified characters (shinjitai), many of which differ from their Chinese counterparts, as well as characters t…

Some.

I know that not every letter has been Han-unified, but I don't know the specifics.

Is it possible, that this problem of yours has actually been taken care of?

If not, is it possible that it's just a mistake instead of a big evil conspiracy?

Re: I Can’t Write My Name in Unicode

#136

Earlier quoted context omitted.

> some languages are insanely complex to implement I don't understand this, is there more to implementing a language than creating glyphs for its character set? I wouldn't think the linguistic complexity would matter at all, only the number of glyphs in the 'alphabet' or similar?

It really depends on the language, and it's not totally about the glyphs. Text entry is a huge challenge for languages like Mandarin where glyphs can have multiple pronunciations and meanings depending on context. Consider that Mandarin (which shares many, but not all glyphs with Japanese kanji) has upwards of 20,000 different glyphs, and that other languages have a similar level of complexity, and it becomes hard to…

User interface isn't the problem, though - bitwise representation is the problem. How do we represent all the valid characters in Unicode? Data entry is an entirely separate issue (as is display).

Re: I Can’t Write My Name in Unicode

#137

> He proudly announces that there are ‘no fewer than 147 Indian dialects’ – a pathetically inaccurate count. Wow. How can a country function like this? Is everyone proficient in their native language plus a 'common' one, or are all interactions supposed to be translated inside the same country? Regardless of historical and cultural value, if that's the case, it seems... inefficient. I do realize that there are more c…

> How can a country function like this? Is everyone proficient in their native language plus a 'common' one

Yes, there may be multiple "common" languages if the country is large/populated enough or for historical reasons (Switzerland has 4 official languages — though official acts only have to be provided in 3 of them, Swiss-German, French and Italian — and they only have 8 million people)

> Regardless of historical and cultural value, if that's the case, it seems... inefficient.

Well telling people to fuck off with their generations-old gobbledygook and imposing a brand new language on them tends to make them kind-of restless, so unless you're willing to assert your declarations in blood (or at least in the specific suppression of non-primary languages)…

The latter has happened quite bit, e.g. post-revolutionary France tried very hard to stamp out both dialects and non-french languages until very recently (the EU shows a distinct lack of appreciation for the finer points of linguicide and linguistic discrimination), the UK did the same throughout the Empire.

> I do realize that there are more countries like this

Almost all of them aside from former british colonies (where native languages were by and large eradicated), at least to an extent (many countries carried out linguistic unification as part of their rise as nation-states, to various levels of success).

Re: I Can’t Write My Name in Unicode

#138
post #89
post #81

Earlier quoted context omitted.

You don't determine language based on codepage. I give you ASCII text; what language is it?

BINGO. But because of Han unification I all of a sudden DO need to know the language. The same Unicode code point needs to be rendered differently for a user in Mainland China versus a user in Japan or else the user may not be able to read the text! Even if the user can read the character, they are going to experience a degradation in reading speed and comprehension, and be generally frustrated. Not to mention showin…

In what situations do you need to do this, but don't need to show any other data (dates and times, localized UI, user timezone, culturally appropriate fonts, RTLness) that involves knowing the user's languages and locale?

This can happen if the user is intentionally reading mixed-language text or text not in their computer's UI language, of course. In that case different CJK languages also have different preferred fonts, so having language tagging or just guessing is pretty important.

Re: I Can’t Write My Name in Unicode

#139
post #82
post #68

Earlier quoted context omitted.

The example characters expressed: 日、中、力 however are written the same way in both Chinese and Japanese from my understanding. (Albeit, I studied Japanese). There are admittedly variations which should be done separately, however unification of visually identical glyphs is a "good thing" imho

I'm fluent in Japanese and speak some Mandarin Chinese as well. These 3 characters are identical, not similar. For a different example, 国 and 國 used to be the same character, but China and Japan (left) have both diverged the traditional form still used in Taiwan (right). Unicode treats them as separate. 今 Looks slightly different in traditional Chinese vs other languages. In traditional Chinese, the little straight l…

This seems again to be a perfect place for rendering rather than encoding. The english letter 'a' can be rendered as a ring with a tail (the way I handwrite), or a ring with a cap and a tail (the way the font usually renders). Both are the same letter, if rendered differently based on my (contextually sensitive) font.

Re: I Can’t Write My Name in Unicode

#140
post #101
post #87

Earlier quoted context omitted.

> The vast majority of Japanese and Chinese characters are not only similar, they are identical. Not all are. Some are clearly different characters deriving from a common historical root, and should not be unified. And what about traditional versus simplified? Which glyph set do I use? Oh wait, thanks to Han unification, I now need to rely on bloody environment variables to decide! For Chinese text, rendering a strin…

Antiqua, Fractura, Schwabacher, Textura and all those other charming variants of writing European languages don't have seperate code points for all those presentational variations. Do we white Western men happen to discriminate against ourselves?

If you don't know the specifics, why the hell are you even arguing this point?
Post reply on HN