Live data from Hacker News

Unicode character “ꙮ” (U+A66E) is being updated

twitter.com

51–60 of 254 posts

Re: Unicode character “ꙮ” (U+A66E) is being updated

#51
post #30

Earlier quoted context omitted.

Honestly it probably deserves the Pluto treatment: decertification as a character. One historical use in the 1400s doesn't merit a character and never did.

Unicode's mission is to make every document "roundtrip-able". Even if a character is only used once, it should be possible to save a plaintext version of the containing document without losing any information. Roughly, I should be able to put a transcription of that one translation from the 1400s on Wikisource without using images. You may disagree with me, and that's fine, but it doesn't change Unicode's mission. Be…

Today, I wrote a document by hand containing a new symbol that only looks like genitalia if you squint really hard. Where do I apply to have it included in unicode so that it can be digitized properly?

Re: Unicode character “ꙮ” (U+A66E) is being updated

#53
post #30
post #4

By the same reasoning, the 7-eyed O has now been used more than once, so it deserves a glyph! So the right way to do this is to introduce a new character for the correct glyph, and also leave the current one (perhaps changing the title). Otherwise these tweets won't make when read by someone that updated to Unicode 15.0

Honestly it probably deserves the Pluto treatment: decertification as a character. One historical use in the 1400s doesn't merit a character and never did.

At the moment this character is used in many documents and databases - including comments in this thread, the article mentioned there, etc.

There could have been a good case not to include it back in 2007, but once it has been included, excluding it would break stuff.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#54
post #35

I don't understand why this character needs to exist given that, at least according to the author, it has only been seen once in the wild, and it's semantically identical to another more widely used character. I'm glad I'm not responsible for unicode. Clearly I have the wrong mindset for it.

Imagine you’re a historian from the future studying some old document, and you spot a weird character that you’ve never seen before. Wouldn’t it be useful to be able to search for that character to see if it shows up in any other document? A simple OCR scan will bring up all the information you could ever need for that one weird symbol.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#55
post #35

I don't understand why this character needs to exist given that, at least according to the author, it has only been seen once in the wild, and it's semantically identical to another more widely used character. I'm glad I'm not responsible for unicode. Clearly I have the wrong mindset for it.

Perhaps it's relevant to look at how it was introduced - as a "package deal" with many, many characters from medieval cyrillic literature, as described in this proposal https://www.unicode.org/L2/L2007/07003r-n3194r-cyrillic.pdf

It certainly made sense to include this package in Unicode, and the vast majority of those characters certainly should be in this proposal. You do have to draw the line somewhere, and obviously those close to the line will be debatable, no matter where you chose to draw it, like this particular symbol - but once you've decided that you will include the one-eyed O (small and capital) and the two-eyed O (small and capital), then putting in the many-eyed O as well to complete the set doesn't seem so far-fetched.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#57
Related thread, about non-existent CJK characters ending up in Unicode through transcription mistakes ("ghost characters"):

https://news.ycombinator.com/item?id=32095502 ("A Spectre Is Haunting Unicode", 180 comments)

edit to add: The top thread in the 2020 repost was about ꙮ,

https://news.ycombinator.com/item?id=24955536

Re: Unicode character “ꙮ” (U+A66E) is being updated

#58
post #47

Earlier quoted context omitted.

Unicode's mission is to make every document "roundtrip-able". Even if a character is only used once, it should be possible to save a plaintext version of the containing document without losing any information. Roughly, I should be able to put a transcription of that one translation from the 1400s on Wikisource without using images. You may disagree with me, and that's fine, but it doesn't change Unicode's mission. Be…

Unicode doesn't have a character for every illuminated initial, nor should it. I'm not clear on why this character should be considered any differently.

Because it's already been added to unicode. Now it's not a question of whether or not to add, rather to remove, and unicode almost by definition does not remove.
Post reply on HN