Live data from Hacker News

Unicode character “ꙮ” (U+A66E) is being updated

twitter.com

171–180 of 254 posts

Re: Unicode character “ꙮ” (U+A66E) is being updated

#171
post #11

When my kids were young, I accidentally flubbed the pronunciation of "Santa Claus" once and said something that sounded a lot like "Centiclops", which I decided to roll with. Centiclops is a lot like a cyclops with one eye, except the as a reading of the roots clearly indicates, this is a creature with 100 eyes. Today I learn that Centiclops effectively has a Unicode character. As Centiclops' representative in the wo…

> Centiclops is a lot like a cyclops with one eye, except th[at] as a reading of the roots clearly indicates, this is a creature with 100 eyes. Not in any normal sense of "roots". Cent is a Latin root meaning 100. ops is a Greek form meaning eye. The -i- indicates that the word is being formed in Latin, and the -cl- is entirely spurious. The original Greek word divides as cycl-ops, not cy-clops.

Cent is easy to grasp if you speak a Romance.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#172

Earlier quoted context omitted.

> I thought "santa" meant "saint"? Well, santa is a Spanish word meaning "holy" and saint is a cognate French word meaning the same thing. They descend from Latin sanctus ; compare sanctify . When the prayer goes "holy Mary, mother of god", "holy Mary" is an exact equivalent of "santa María".

Might as well mention “Sancta Marīa” in Latin, for example from the Christian Hail Mary[1], a recorded Latin version[2], written Latin next to English and Spanish[3] and of course translated into thousands of languages[4] although unfortunately mostly written using /A-Z/i ; I am an atheist interested in languages. [1] https://en.m.wikipedia.org/wiki/Hail_Mary [2] https://glaemscrafu.jrrvf.com/english/avemaria.html [3…

In my mind, the Latin form of Mary is Mariam, because that's what my Latin teacher taught me. (He also commented that, unlike Greek names, Hebrew names never inflected in Latin, so that it would be "Mariam" regardless of what case the name should appear in.)

But it makes sense that Church Latin would be different.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#173
post #171

Earlier quoted context omitted.

> Centiclops is a lot like a cyclops with one eye, except th[at] as a reading of the roots clearly indicates, this is a creature with 100 eyes. Not in any normal sense of "roots". Cent is a Latin root meaning 100. ops is a Greek form meaning eye. The -i- indicates that the word is being formed in Latin, and the -cl- is entirely spurious. The original Greek word divides as cycl-ops, not cy-clops.

Cent is easy to grasp if you speak a Romance.

But it doesn't combine with ops. You'd need to talk about a hecatops or a hecatontops. And even more than it can't combine with ops, it can't combine with clops because there is no such root.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#174
post #122

Earlier quoted context omitted.

If that was once its mission, it was clearly abandoned long ago. They rejected Klingon characters on the grounds that it has low usage for communication, and that many of the people who do communicate in Klingon use a latinized form. ꙮ seems to just be a fancy way of writing О. I haven't seen anything that says it has a different meaning. The arguments for excluding Klingon seem to apply even more so to ꙮ.

If you look through the old mailing list postings, the oft-left-implicit problem with Klingon (as well as Tengwar, Everson’s [EDIT: misspelling] pet project) is that it may get people into legal trouble (even though in a reasonable world it shouldn’t be able to). So in the unofficial CSUR / UCSUR they remain. A weird solitary character from the 1400s isn’t subject to that, and even if it’s a mistake it’s probably not…

[EDIT: Removed a section about the now-fixed typo]

> I think the last such break with code points genuinely changing meanings was to repair a mistaken CJK unification some time in the 00s, and the Consortium may even have tied its own hands in that regard with the ever-more-strict stability policies[.]

Not exactly, the last break happened between Unicode 1.1 and 2.0 and the new CJK Unified Ideographs Extension A block still contains unified characters. The main reason for break was that both Hangul and CJK(V) ideographs required tons of additional code points and it became clear that 16-bit code space is dangerously insufficient; by 1.1 there was only a single big block of unassigned code points from U+A000 to U+E7FF (18,432 total), and there were 4,516 and 6,582 new Hangul and CJK(V) ideographs in 2.0 (11,098 total).

Re: Unicode character “ꙮ” (U+A66E) is being updated

#175
post #30
post #4

By the same reasoning, the 7-eyed O has now been used more than once, so it deserves a glyph! So the right way to do this is to introduce a new character for the correct glyph, and also leave the current one (perhaps changing the title). Otherwise these tweets won't make when read by someone that updated to Unicode 15.0

Honestly it probably deserves the Pluto treatment: decertification as a character. One historical use in the 1400s doesn't merit a character and never did.

There are characters in unicode with 0 usages that we dont even know where they came from. E.g. 彁

Re: Unicode character “ꙮ” (U+A66E) is being updated

#176
post #51

Earlier quoted context omitted.

Unicode's mission is to make every document "roundtrip-able". Even if a character is only used once, it should be possible to save a plaintext version of the containing document without losing any information. Roughly, I should be able to put a transcription of that one translation from the 1400s on Wikisource without using images. You may disagree with me, and that's fine, but it doesn't change Unicode's mission. Be…

Today, I wrote a document by hand containing a new symbol that only looks like genitalia if you squint really hard. Where do I apply to have it included in unicode so that it can be digitized properly?

Given that “𓂸” (U+130B8) is already in unicode (and related 𓂹,𓂺) pretty sure the only problem is you made it up, not that it looks like genitilia

Re: Unicode character “ꙮ” (U+A66E) is being updated

#177

Earlier quoted context omitted.

One could argue that emoji should have never been added to Unicode in the first place. Peaches and butts are images, pictures, illustrations, whatever - but they are not characters. There's no writing system which has a colored drawing of a peach as a character.

But that doesn't change the fact that most people use them snd like them, and there is not much technical disruption. They just chose practicality over purity.

Not only that - people use them in textual communication the way letters traditionally are used. There is probably a lot better argument for emoiji than a lot of other things in unicode (but it is a slippery slope)

Re: Unicode character “ꙮ” (U+A66E) is being updated

#178
post #11

When my kids were young, I accidentally flubbed the pronunciation of "Santa Claus" once and said something that sounded a lot like "Centiclops", which I decided to roll with. Centiclops is a lot like a cyclops with one eye, except the as a reading of the roots clearly indicates, this is a creature with 100 eyes. Today I learn that Centiclops effectively has a Unicode character. As Centiclops' representative in the wo…

> Centiclops is a lot like a cyclops with one eye, except th[at] as a reading of the roots clearly indicates, this is a creature with 100 eyes. Not in any normal sense of "roots". Cent is a Latin root meaning 100. ops is a Greek form meaning eye. The -i- indicates that the word is being formed in Latin, and the -cl- is entirely spurious. The original Greek word divides as cycl-ops, not cy-clops.

Impressively polylingual, even multiglot.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#179
post #30

Earlier quoted context omitted.

Honestly it probably deserves the Pluto treatment: decertification as a character. One historical use in the 1400s doesn't merit a character and never did.

There are characters in unicode with 0 usages that we dont even know where they came from. E.g. 彁

While the origin of 彁 will never be certain, there is a good chance that it came from a misinterpretation of 彊 [1]. Why is this not an accepted theory though? Because it is still possible that 彁 did appear in some reference source from the standardization, and neither that source or a source where 彊 does look like 彁 was found.

[1] http://www.asahi.com/special/kotoba/archive2015/moji/2011082...

Re: Unicode character “ꙮ” (U+A66E) is being updated

#180
post #75

Earlier quoted context omitted.

I thought "santa" meant "saint"?

“Santa” means “female saint” in Italian and Spanish. Perhaps the English “santa” came from another language but I always found the name “Santa Claus” just horrible.

The Tim Allen movie series in Spanish is titled "Santa Cláusula".
Post reply on HN