Live data from Hacker News

Unicode character “ꙮ” (U+A66E) is being updated

twitter.com

161–170 of 254 posts

Re: Unicode character “ꙮ” (U+A66E) is being updated

#161
post #122

Earlier quoted context omitted.

Unicode's mission is to make every document "roundtrip-able". Even if a character is only used once, it should be possible to save a plaintext version of the containing document without losing any information. Roughly, I should be able to put a transcription of that one translation from the 1400s on Wikisource without using images. You may disagree with me, and that's fine, but it doesn't change Unicode's mission. Be…

If that was once its mission, it was clearly abandoned long ago. They rejected Klingon characters on the grounds that it has low usage for communication, and that many of the people who do communicate in Klingon use a latinized form. ꙮ seems to just be a fancy way of writing О. I haven't seen anything that says it has a different meaning. The arguments for excluding Klingon seem to apply even more so to ꙮ.

If you look through the old mailing list postings, the oft-left-implicit problem with Klingon (as well as Tengwar, Everson’s [EDIT: misspelling] pet project) is that it may get people into legal trouble (even though in a reasonable world it shouldn’t be able to). So in the unofficial CSUR / UCSUR they remain.

A weird solitary character from the 1400s isn’t subject to that, and even if it’s a mistake it’s probably not worth breaking compatibility at this point (I think the last such break with code points genuinely changing meanings was to repair a mistaken CJK unification some time in the 00s, and the Consortium may even have tied its own hands in that regard with the ever-more-strict stability policies).

Similarly, for example, old ISO keyboard symbols (the ⌫ for erase backwards, but also a ton of virtually unused ones) were thrown in indiscriminately at the beginning of the project when attempting to cover every existing encoding, but when the ISO decided to extend the repertoire they were told to kindly provide examples of running-text (not iconic) usage in a non-member-body-controlled publication. (Crickets. The ISO keyboard input model itself only vaguely corresponds to how input methods for QWERTY-adjacent keyboards work in existing systems—as an attempt at rationalization, it seems to mostly be a failed one.)

Re: Unicode character “ꙮ” (U+A66E) is being updated

#162
post #30

Earlier quoted context omitted.

Honestly it probably deserves the Pluto treatment: decertification as a character. One historical use in the 1400s doesn't merit a character and never did.

One historical use in the 1400s doesn't merit a character and never did One known and surviving use. It is possible that it exists in other places, since the vast majority of the planet's written work has not been digitized. It may also have been used other places that have not survived. Just because it's not important to you does not mean it is not important. The fact that is survived for 600 years makes it interest…

The thing is, looking at the page, there are many other characters that were not added - the large red С-looking characters, for example. But for some "bizarre" reason, those were not included in Unicode...

Of course, the simple answer is that Unicode actually includes any character that someone cares enough to ask to be added, with rare exceptions.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#163

Here[1][2] is the scan of manuscript from 1429, image #251 [1] https://lib-fond.ru/lib-rgb/304-i/f-304i-308/#image-251 [2] https://web.archive.org/web/20110927102700/https://www.stsl....

So the text at that point literally talks about ‘many-eyes seraphims’. The eyes symbol is a pure gag—seems to be spliced in place of the letter ‘о’ in the word ‘eye’ just a little down the line. (However, Old Slavonic is a tough read due to no spaces, so I'm not sure about that word. But at least it's not the Glagolitic script, which was just ridiculous and actually had multi-circle letters.)

Re: Unicode character “ꙮ” (U+A66E) is being updated

#164
post #66
post #47

Earlier quoted context omitted.

Unicode doesn't have a character for every illuminated initial, nor should it. I'm not clear on why this character should be considered any differently.

http://std.dkuug.dk/jtc1/sc2/wg2/docs/n3194.pdf It was introduced with other "ocular O"s which are seemingly more commonly used than this one. It's not quite an illuminated initial.

Wow, this is probably the most actually useful and interesting comment in this whole discussion, thanks! For anyone interested, the most relevant quotes from the document are in particular:

"This document requests the addition of a number of Cyrillic characters to be added to the UCS. It also requests clarification in the Unicode Standard of four existing characters. This is a large proposal. While all of the characters are either Cyrillic characters (plus a couple which are used with the Cyrillic script), they are used by different communities. Some are used for non-Slavic minority languages and others are used for early Slavic philology and linguistics, while others are used in more recent ecclesiastical contexts. We considered the possibility of dividing the proposal into several proposals, but since this proposal involves changes to glyphs in the main Cyrillic block, adds a character to the main Cyrillic block, adds 16 characters to the Cyrillic Supplement block, adds 10 characters to the new Cyrillic Extended-A block currently under ballot, creates two entirely new Cyrillic blocks with 55 and 26 characters respectively, as well as adding two characters to the Supplementary Punctuation block, it seemed best for reviewers to keep everything together in one document.

(...)

MONOCULAR O Ꙩꙩ, BINOCULAR O Ꙫꙫ, DOUBLE MONOCULAR O Ꙭꙭ, and MULTIOCULAR O ꙮ are used in words which are based on the root for ‘eye’. The first is used when the wordform is singular, as ꙩкꙩ; the second and third are used in the root for ‘eye’ when the wordform is dual, as ꙫчи, ꙭчи; and the last in the epithet ‘many-eyed’ as in серафими многоꙮчитїй ‘many-eyed seraphim’. It has no upper-case form. See Figures 34, 41, 42, 55."

Re: Unicode character “ꙮ” (U+A66E) is being updated

#165
post #122

Earlier quoted context omitted.

Unicode's mission is to make every document "roundtrip-able". Even if a character is only used once, it should be possible to save a plaintext version of the containing document without losing any information. Roughly, I should be able to put a transcription of that one translation from the 1400s on Wikisource without using images. You may disagree with me, and that's fine, but it doesn't change Unicode's mission. Be…

If that was once its mission, it was clearly abandoned long ago. They rejected Klingon characters on the grounds that it has low usage for communication, and that many of the people who do communicate in Klingon use a latinized form. ꙮ seems to just be a fancy way of writing О. I haven't seen anything that says it has a different meaning. The arguments for excluding Klingon seem to apply even more so to ꙮ.

Unless it's legitimately someone's native tongue, conlangs shouldn't be in unicode. If there are kids out there that are native Klingon speakers, then you can make the argument it should be included.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#166

Earlier quoted context omitted.

It could go either way, it is not always that the scientific meaning wins out, especially not when even scientists don’t find the new definition useful. When I think of a planet, I think of a world that has active geology that isn’t a moon (I know excluding moons is arbitrary, and perhaps I shouldn’t do that; but hey, that’s language for you). I honestly don’t care about the orbit, and I bet that when most people thi…

> When I think of a planet, I think of a world that has active geology Wouldn't that definition rule out gas giants?

No just that, but whether or not Mars is still geologically active is still an open question. If you admit planets on the basis that they have a history of geological activity, then Ceres is a planet too.

I don’t think anybody considers geological activity as particularly useful for classifying things as ‘planet’ or ‘not planet’.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#167

Earlier quoted context omitted.

> How much work would it be to implement your own font of the entire unicode set? Or is that not actually a thing and fonts implement as-desired subsets? You can't, and you are not expected to do so. You are limited by OpenType limit (65,535 glyphs), various shaping rules that possibly increase the number of required glyphs, and lack of local or historical typographic convention. Your best bet is either to recruit a…

A single OpenType font file is limited to 65,535 glyphs. Nothing stops your font from being implemented as a series of .otf files (besides what people think of as a "font" when it comes to usage on computers). But yes, time constraints are the limiting factor. I don't think anyone is going to dedicate their entire life to making a single font.

While you are right that one logical font can consist of multiple font files (or possibly a OpenType collection), this constraint does affect most typical fonts, and in particular wide-coverage CJK fonts already hit this limit. Fonts supporting only one of Chinese, Japanese and Korean don't need that many glyphs, and probably even two of them will be okay, but fonts with all three sets of glyphs won't. It is therefore common to provide three versions of fonts, all differently named.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#168
post #4

By the same reasoning, the 7-eyed O has now been used more than once, so it deserves a glyph! So the right way to do this is to introduce a new character for the correct glyph, and also leave the current one (perhaps changing the title). Otherwise these tweets won't make when read by someone that updated to Unicode 15.0

why not make an additional eye a diacritic mark so you can just add an arbitrary number of eyes

Re: Unicode character “ꙮ” (U+A66E) is being updated

#169

Earlier quoted context omitted.

> When I think of a planet, I think of a world that has active geology Wouldn't that definition rule out gas giants?

No just that, but whether or not Mars is still geologically active is still an open question. If you admit planets on the basis that they have a history of geological activity, then Ceres is a planet too. I don’t think anybody considers geological activity as particularly useful for classifying things as ‘planet’ or ‘not planet’.

Why shouldn’t Ceres be a planet? If Pluto gets to be a planet then Ceres is definitely a planet.

But there is still active geology on Mars. There is still moisture, winds and ice-caps that are shaping the environment. I consider that to be geologically active.

EDIT: And there are actual experts which consider active geology (or something similar) to be a planet, including Anton Petrov (https://www.youtube.com/watch?v=8-2HxrgqUnM)

Re: Unicode character “ꙮ” (U+A66E) is being updated

#170

Here[1][2] is the scan of manuscript from 1429, image #251 [1] https://lib-fond.ru/lib-rgb/304-i/f-304i-308/#image-251 [2] https://web.archive.org/web/20110927102700/https://www.stsl....

Looks more like a diagram in the middle of text. It's very unique. It should not be a character
Post reply on HN