Live data from Hacker News

Unicode character “ꙮ” (U+A66E) is being updated

twitter.com

111–120 of 254 posts

Re: Unicode character “ꙮ” (U+A66E) is being updated

#111

Here[1][2] is the scan of manuscript from 1429, image #251 [1] https://lib-fond.ru/lib-rgb/304-i/f-304i-308/#image-251 [2] https://web.archive.org/web/20110927102700/https://www.stsl....

It's curious that the red ink blobs behind the "eyes" aren't included in the unicode glyph either...

Re: Unicode character “ꙮ” (U+A66E) is being updated

#112

Quoted post unavailable.

I guess because the goal of Unicode is to be able to represent every character that's appeared in language. This one is in a published book, while guns and a sexual intercourse symbol aren't. Emoji was a weird value add that Japanese mobile providers added to their phones before Unicode. To get them to move to Unicode, they had to keep them. That's why there's a Tokyo Tower emoji, but not an Eiffel Tower. That's why…

That seems actually logical when you consider that kanji presumably began as simple depictions of objects that could be drawn quickly. Perhaps the only difference between emoji and kanji is time.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#113
post #4

By the same reasoning, the 7-eyed O has now been used more than once, so it deserves a glyph! So the right way to do this is to introduce a new character for the correct glyph, and also leave the current one (perhaps changing the title). Otherwise these tweets won't make when read by someone that updated to Unicode 15.0

This thread on HN won't make sense in the future if the Unicode body replaces ꙮ

Make a new character!

Re: Unicode character “ꙮ” (U+A66E) is being updated

#114
post #14
post #8

Earlier quoted context omitted.

Unicode basic rule is that character definitions never ever change, even when enumerated erroneously.

Yes, but this is a change either way, because that codepoint's definition referred to that character. Either the reference or the description of the appearance has to change.

   ꙮ ꙮ 
  ꙮ ꙮ ꙮ 
   ꙮ ꙮ 
ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ

Make a new character. Updating the existing character ruins the meaning of all previous usages.

It's like trying to change an API. Don't disrespect your existing users. Make a new version.

(ꙮ ͜ʖꙮ)

Think of all the ASCII art this botches. That has to have some historical importance to the Unicode standards body.

(⌐ꙮ_ꙮ)

For scholarly digital (unprinted) documents where the correct character rendering matters, erroneous past usages can be trivially found with grep, a date search, and easily corrected. The domain experts will familiarize themselves with this issue and fix the problem. Don't take a shotgun to it!

This message wꙮn't have the ꙮriginally intended meaning if the characters are updated from underneath.

ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ

Re: Unicode character “ꙮ” (U+A66E) is being updated

#115

Earlier quoted context omitted.

For as inclusive as that mission is, it seems weird to me how limited in certain areas unicode is. For instance, people use peach emoji since there isn't one for butt, eggplant since there's no penis, etc. This doesn't contradict the stated goal exactly, but it seems against the spirit of it at least.

One could argue that emoji should have never been added to Unicode in the first place. Peaches and butts are images, pictures, illustrations, whatever - but they are not characters. There's no writing system which has a colored drawing of a peach as a character.

[deleted]

Re: Unicode character “ꙮ” (U+A66E) is being updated

#116
post #75

Earlier quoted context omitted.

“Santa Clause” would translate to “holy clause”. There might be such a thing but I think you meant Santa Claus :)

I thought "santa" meant "saint"?

I thought it was a misspelling of Satan, but maybe that's because I'm Jewish.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#117

Earlier quoted context omitted.

Uff. I'm not sure we have space for another glyph in Unicode. Looks pretty packed in here...

UTF-8 is still more than 80% empty, and can be potentially extended...

Theoretically, UTF-8 can encode up to 31 bits (U+7FFF'FFFF)[0], but for compatibility with UTF-16's surrogates, it's officially capped to 21 bits with the max being U+10'FFFF[1]. That decision was made November 2003, so there's two decades of software written with hard caps of U+10'FFFF.

[0]: https://www.rfc-editor.org/rfc/rfc2279

[1]: https://www.rfc-editor.org/rfc/rfc3629#section-3

Re: Unicode character “ꙮ” (U+A66E) is being updated

#119
post #35

I don't understand why this character needs to exist given that, at least according to the author, it has only been seen once in the wild, and it's semantically identical to another more widely used character. I'm glad I'm not responsible for unicode. Clearly I have the wrong mindset for it.

Surprisingly many characters in Unicode are only recorded a few times if not once before the assignment. Chinese characters for example have a lot of them, because it was relatively frequent to make a new character for newborns before the modernity and some of them have survived through literatures but otherwise seen no uses (e.g. 𡸫 U+21E2B only appears once in the Records of the Three Kingdoms 三國志). But they have still received code points because they are considered essential for digitaization of historical works, and multiocular O is no different.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#120

Wait a minute, how will we refer to the old glyph in the future? Once this is updated the articles such as this one will have the new shape.

There was a joke that U+A66E should retain seven eyes and further eyes should be added with a ZWJ sequence [1]. If that character somehow got very popular in modern texts, updating its glyph may result in an interoperability problem so such solution would have been needed. But that didn't happen so the glyph itself has been updated instead.

[1] https://twitter.com/BabelStone/status/1323440365429542919

Post reply on HN