Live data from Hacker News

Unicode character “ꙮ” (U+A66E) is being updated

twitter.com

181–190 of 254 posts

Re: Unicode character “ꙮ” (U+A66E) is being updated

#181

Earlier quoted context omitted.

For as inclusive as that mission is, it seems weird to me how limited in certain areas unicode is. For instance, people use peach emoji since there isn't one for butt, eggplant since there's no penis, etc. This doesn't contradict the stated goal exactly, but it seems against the spirit of it at least.

One could argue that emoji should have never been added to Unicode in the first place. Peaches and butts are images, pictures, illustrations, whatever - but they are not characters. There's no writing system which has a colored drawing of a peach as a character.

They're sort of neither. The peach emoji will render differently on iOS, Android, Windows. And I'm sure emoji-replacement packs are possible on Windows and Android (even though it's also guaranteed to be a virus).

So a peach emoji is not the same thing as the iOS peach-emoji-image. Similar to how changing my font doesn't change the actual characters.

I don't think including emojis was a great idea, but now that it's happened and people everywhere use them, emoji have become characters. I agree with your point, but it's already happened and so now there's not really any going back.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#182

Here[1][2] is the scan of manuscript from 1429, image #251 [1] https://lib-fond.ru/lib-rgb/304-i/f-304i-308/#image-251 [2] https://web.archive.org/web/20110927102700/https://www.stsl....

Looks more like a diagram in the middle of text. It's very unique. It should not be a character

It's used in place of the letter "o", so not purely a diagram but it feels like the role of a font to me, not a dedicated character.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#183

I’m not sure how I feel about this. I’m not an expert by any means. But something just doesn’t feel right when you’ve got unicode with a character with one known use from forever ago. Doesn’t this open up the flood gates to just a ridiculous amount of work or else biased gatekeeping? How much work would it be to implement your own font of the entire unicode set? Or is that not actually a thing and fonts implement as-…

> How much work would it be to implement your own font of the entire unicode set? Or is that not actually a thing and fonts implement as-desired subsets? You can't, and you are not expected to do so. You are limited by OpenType limit (65,535 glyphs), various shaping rules that possibly increase the number of required glyphs, and lack of local or historical typographic convention. Your best bet is either to recruit a…

You could also go the shady route and just make a font out of all the "reference character sheets" that the Unicode site has. Probably not legal and the result would not be pleasant to read, but that's one way to create a font containing all of Unicode.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#184
post #30

Earlier quoted context omitted.

Honestly it probably deserves the Pluto treatment: decertification as a character. One historical use in the 1400s doesn't merit a character and never did.

Unicode's mission is to make every document "roundtrip-able". Even if a character is only used once, it should be possible to save a plaintext version of the containing document without losing any information. Roughly, I should be able to put a transcription of that one translation from the 1400s on Wikisource without using images. You may disagree with me, and that's fine, but it doesn't change Unicode's mission. Be…

The artist Prince changed his stage name to an unpronounceable symbol for a few years. It appears in more than one document. Should it be added to Unicode?

Re: Unicode character “ꙮ” (U+A66E) is being updated

#186

Earlier quoted context omitted.

“Santa” means “female saint” in Italian and Spanish. Perhaps the English “santa” came from another language but I always found the name “Santa Claus” just horrible.

The name Santa Claus evolved from Nick's Dutch nickname, Sinter Klaas, a shortened form of Sint Nikolaas (Dutch for Saint Nicholas) https://www.history.com/.amp/topics/christmas/santa-claus#si...

It’s actually Sinterklaas (without a space) and we still call him that :) We also ended up re-importing the American Santa Claus, so these days we have two festive holidays in December.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#187
post #158

Earlier quoted context omitted.

Unicode's mission is to make every document "roundtrip-able". Even if a character is only used once, it should be possible to save a plaintext version of the containing document without losing any information. Roughly, I should be able to put a transcription of that one translation from the 1400s on Wikisource without using images. You may disagree with me, and that's fine, but it doesn't change Unicode's mission. Be…

Meanwhile one still can't roundtrip regular Japanese without some kind of funky out-of-band signalling. By itself this kind of thing is harmless, but it speaks to poor prioritization from Unicode.

Why can’t it round-trip Japanese?

Re: Unicode character “ꙮ” (U+A66E) is being updated

#188
post #122

Earlier quoted context omitted.

If that was once its mission, it was clearly abandoned long ago. They rejected Klingon characters on the grounds that it has low usage for communication, and that many of the people who do communicate in Klingon use a latinized form. ꙮ seems to just be a fancy way of writing О. I haven't seen anything that says it has a different meaning. The arguments for excluding Klingon seem to apply even more so to ꙮ.

Unless it's legitimately someone's native tongue, conlangs shouldn't be in unicode. If there are kids out there that are native Klingon speakers, then you can make the argument it should be included.

I think it makes way more sense to put a conlang in Unicode than it does a peculiar stylistic flourish only ever applied once to a single letter in a single document. If that belongs in Unicode, why not every bit of marginalia ever doodled and every uniquely adorned drop cap / initial letter?

Re: Unicode character “ꙮ” (U+A66E) is being updated

#189

Earlier quoted context omitted.

> Centiclops is a lot like a cyclops with one eye, except th[at] as a reading of the roots clearly indicates, this is a creature with 100 eyes. Not in any normal sense of "roots". Cent is a Latin root meaning 100. ops is a Greek form meaning eye. The -i- indicates that the word is being formed in Latin, and the -cl- is entirely spurious. The original Greek word divides as cycl-ops, not cy-clops.

A bit like the heli-copter | helico-pter thing.

Wow. I never noticed that before. Spiral wings.

Re: Unicode character “ꙮ” (U+A66E) is being updated

#190
post #158

Earlier quoted context omitted.

Meanwhile one still can't roundtrip regular Japanese without some kind of funky out-of-band signalling. By itself this kind of thing is harmless, but it speaks to poor prioritization from Unicode.

Why can’t it round-trip Japanese?

"Han Unification" - in Unicode many Japanese characters are represented as Chinese characters that look different (and subjectively ugly). The Unicode consortium's answer is that you're supposed to use a different font or something when displaying Japanese, which is pretty unsatisfying (e.g. if you want to have a block of text that contains both Japanese and Chinese, you can't represent that as just a Unicode string, it has to be some kind of rope of segments with their own fonts, at which point frankly you might as well just go back to bytes-with-encoding which at least breaks very clearly and visibly if you get it wrong).
Post reply on HN