Live data from Hacker News

How a comment on Hacker News led to 4½ new Unicode characters

unicodepowersymbol.com

391–400 of 429 posts

Re: How a comment on Hacker News led to 4½ new Unicode characters

#391
post #267

Earlier quoted context omitted.

> I can have a document which combines English, Russian, Arabic, and Chinese English combines with top-down Chinese, and English combines with right-to-left Arabic, but top-down Chinese and right-to-left Arabic don't combine properly in the same document using Unicode -- the Arabic will be written bottom-up instead of top-down when embedded in the top-down Chinese.

Isn't that not not a problem with Unicode, but the text rendering engine? Or is this indeed a spec bug?

Right-to-left script rendering is part of the Unicode spec, whereas top-down scripting for Chinese is a font issue.

Re: How a comment on Hacker News led to 4½ new Unicode characters

#392

Earlier quoted context omitted.

We could also use "OHM" instead of Ω and "EUR" instead of €, but symbols provide more concise representations of meaning, especially in the case of a combined ON/OFF button. Much easier to fit the universal circle with 1 in it than "ON/OFF". The broken circle with 1 in it is way smaller than "ON/STANDBY". The other two symbols are then necessary to keep the proportions consistent across all four related glyphs.

That's the only argument I've seen for it that makes any sense. I'll counter by saying only "ON" is needed, not "ON/OFF", as the off part is implicit. Same goes for STANDBY. This has been standard practice for a long time, I'm not just making things up on the fly. BTW, SBY is a standard abbreviation for STANDBY used by the military, if space is a problem. And SBY is a lot more google-able than an icon.

Yeah, all true. Especially the Google-ability of it.

Ok, here's my last argument then. It's aesthetically pleasing. I admit that's probably the weakest argument, but also the hardest to refute :)

Re: How a comment on Hacker News led to 4½ new Unicode characters

#393
post #362

Earlier quoted context omitted.

How about usage in this passage: 囍事 http://www.chinatimes.com/newspapers/20160623000760-260115 In response to your edit, I mean that it has cultural resonance that is unusually strong in relation to its linguistic overtones, in many ways similar to the semantic timbre of a character like 福. The level of abstraction is different from English because of the ideographic nature of characters that means the visual appeara…

Assuming that 囍事 in that passage refers to "a wedding", first I'd admit that that passes pretty much any test of "linguistic use in running text". Having said that, I note that 喜事 appears in my dictionaries with the gloss "wedding" (well, "any occasion meriting joy, particularly a wedding"), 囍事 does not, and since 囍 is a symbol of weddings which is generally assumed by the Chinese to have the same pronunciation as 喜…

You just aren't getting it. Do you actually know Chinese, or are you just looking things up in a dictionary?

It is pretty natural to jump from 喜 to 囍 because that is how Chinese works. You take radicals, and you bundle them up. You have the "busho" system where people in the past bundled up little bits and pieces and form new words. No reason why people in the present can't do the same.

Re: Micro$oft being outrageous if $ becomes a part of the alphabet. You are misapplying an English oriented viewpoint. In Chinese, there is no objection to forming words in that way, by incorporating radicals together. It's similar in theme to how in German, you can just keep stringing words together to form larger words. In fact, I actually think in the future, words like Micro$oft should entitle $ to become part of the alphabet! That's a very Chinese way of looking at things.

Language is not static. Systems that try to encode language are descriptive. They can never be prescriptive - otherwise we as a civilization die.

If 囍 wants to be a character point, let it be one. If 福倒 wants to be one, there should be one. Isn't the point of unicode to have enough space to include all these kinds of language artifacts (artifact as in a cultural / historical item thought up by humans) in order so people can uniquely reference each one? They are distinct logical units.

If the unicode rulebooks are too rigid, the rules need to change or the approach needs to change. It's useless to try to argue that xyz character in another language shouldn't/can't be a character - people will just stop using unicode if it doesn't suit their needs.

Reeks of colonialism, that's what it is.

EDIT: as an additional gloss, here's why I think 喜 and 囍 are sometimes used differently, even though by the dictionary definition they seem to be the same. I will explain why I think logically they are different concepts.

喜 is happiness, delight, joy. It is probably an adjective in the English sense (I can't map grammar rules through different languages easily).

事 is an occurrence, an item, something that happens.

When you put them together,

喜事 literally means something happy is happening.

The cultural meaning has turned that into a connotation of "wedding", but it could actually be a ton of happy things. Promotions, and yes - one other really big thing in a person's life: having a baby.

有喜 (means "having happiness") is the traditional way of referring to a woman being pregnant

http://baike.baidu.com/view/301829.htm

You can turn that into 家有喜事 - meaning home having something happy - as in this household is having a baby. And you can use it without the 有 - and just use 喜事 to refer to having a baby.

This is different from a wedding.

囍 is a modification of 喜, by doubling up the character and treating it as a radical, people are referring to the idea that there are "two people having happiness" - like a doubled amount of happiness.

In the article linked http://www.chinatimes.com/newspapers/20160623000760-260115 - the 囍事 is used to specifically identify the "wedding" type of 喜事 - it's like trying to avoid the ambiguity and double-entendres that Chinese writing typically embraces and just presents things matter of fact, which is ideal because the article is a newspaper article about customs of towns. Not really something you want people to have multiple interpretations like an essay or a poem, for example.

So logically, there is a difference when trying to use 喜 vs 囍 and I actually really appreciate the author's use of the double version in the text.

I know that not everyone reads these characters in this way, but I do - and I'm sure other people will notice this too. It's the best part of Chinese - not knowing, and not seeing the ambiguity, and one day, someone tells you about it .. and you're like - OMG that's what that means ...

For my earlier indication that this type of character modification is common in chinese:

木 = wood

林 = common last name Lin, also means forest (uncommon on its own)

森 = common character for forest.

The English word "forest" is usually 森林

It's just a doubling and trippling of the 木 radical.

What does it matter that this character is super old - people thousands of years ago thought this up.

Also, if this character weren't so old, would you say that 森 and 林 are both forests and thus don't need separate character points in unicode? That's outrageous!

So now we have a modern version of this modification 喜 -> 囍

And I showed how I think they are different logical concepts.

The link http://www.chinatimes.com/newspapers/20160623000760-260115 showed how it can be used in typographical context.

hmm what's the issue with it being a unicode character point?

POST Edit

In the writing of this post, I think I've come to identify Chinese as an "ambiguity-first" language - I learned Chinese as my mother tongue, but stopped at a elementary school level, and switch over to learning English to a Bachelor's degree level.

In Chinese, puns, double-entendres, and ambiguity just "happens" by default, and you have to work your way to be crystal clear.

English is more straight-forward, with a speaker having to try to make puns or double-entendres.

In the case of 囍, it's a reduction in scope. Modern Chinese people had to create a new word just to narrow down the meaning of 喜 - so that it specifically refers to weddings.

Your whole line of thinking was that 喜 already had meanings inclusive of wedding, so 囍 can't possibly add any more meaning when it also means wedding. In actuality, it took away a bunch of extraneous connotations, and in Chinese, the reduction in complexity is so valuable that it's worth a new word.

I think that it's a mistake to try to over-literate and reduce languages into a set of rulebooks for character encoding - that's all I am trying to put forth - it's best for the person or peoples who speak the language to come up with the encoding for it. I have an elementary school knowledge of Chinese and already I am kinda miffed at why people have an objection to 喜 vs 囍

Imagine how the people who have Bachelor's degrees in Chinese must feel.

Re: How a comment on Hacker News led to 4½ new Unicode characters

#394
post #375

Earlier quoted context omitted.

Seriously, you're able to Google up those other links, but you somehow can't find the (non-exhaustive) Unsupported Scripts list or the Proposed New Scripts pages on the Unicode site? And, without knowing the situation for any of them, you're going to throw out excuses for why the absences don't matter? These aren't characters, but entire scripts that are not part of the standard. Nor are major scripts like kanji comp…

Again, you're the one making the claim. Can you precisely state what you believe to be the problem and cite some sources that this is a major problem and that nobody is working on it? More importantly, ask why it seems unreasonable that a small number of very widely-used ISO standard symbols were incorporated quickly? Wouldn't that be the most reasonable expectation since it lacks the political heat of e.g. Han unifi…

I've already stated my complaint: that getting gratuitous icons into Unicode is easier than actual scripts for human languages. Since you're having Google issues, I'll link you to a page I already mentioned, which it self links to other relevant pages http://unicode.org/standard/unsupported.html

I love the echoing nature of these counter-arguments, that a problem doesn't even exist unless it's "major" and "nobody is working on it". I wonder how many actual different human beings have responded to me in this thread...

Re: How a comment on Hacker News led to 4½ new Unicode characters

#396
post #136

Earlier quoted context omitted.

Semantically, yes. In the code tables you'll see that that "opposite" symbols have links to each other. Programatically, it is much easier to say "does a character lie between 0x12 and 0xBC" than to create a function like `isSymbolForTrafficInEurope()`

Perhaps Unicode could just have tables listing all relevant sequences of symbols, instead. So "latin letters lowercase" would list the codepoints for a-z in order, for example. Would no longer matter if the codepoints themselves are sequential or not. (And relevant to my country, "Swedish characters lowercase" would map to latin letters lowercase + åäö.)

Characters have a script associated with them (e.g. Latin), and caseness is also part of a character's properties.

Now, language-specific subsets¹ of those are a bit iffy to deal with. Especially when text can contain loan words from other languages, so in my experience it's rarely a useful thing to ask for.

¹ Yes, subsets. Latin letters lowercase is not the set abcdefghijklmnopqrstuvwxyz. It is the set

abcdefghijklmnopqrstuvwxyzªºßàáâãäåæçèéêëìíîïðñòóôõöøùúûüýþÿ āăąćĉċčďđēĕėęěĝğġģĥħĩīĭįıijĵķĸĺļľŀłńņňʼnŋōŏőœŕŗřśŝşšţťŧũūŭůűųŵ ŷźżžſƀƃƅƈƌƍƒƕƙƚƛƞơƣƥƨƪƫƭưƴƶƹƺƽƾƿdžljnjǎǐǒǔǖǘǚǜǝǟǡǣǥǧǩǫǭǯǰdzǵǹǻǽǿ ȁȃȅȇȉȋȍȏȑȓȕȗșțȝȟȡȣȥȧȩȫȭȯȱȳȴȵȶȷȸȹȼȿɀɂɇɉɋɍɏɐɑɒɓɔɕɖɗɘəɚɛɜɝɞɟɠɡɢ ɣɤɥɦɧɨɩɪɫɬɭɮɯɰɱɲɳɴɵɶɷɸɹɺɻɼɽɾɿʀʁʂʃʄʅʆʇʈʉʊʋʌʍʎʏʐʑʒʓʕʖʗʘʙʚʛʜʝʞʟ ʠʡʢʣʤʥʦʧʨʩʪʫʬʭʮʯʰʱʲʳʴʵʶʷʸˠˡˢˣˤᴀᴁᴂᴃᴄᴅᴆᴇᴈᴉᴊᴋᴌᴍᴎᴏᴐᴑᴒᴓᴔᴕᴖᴗᴘᴙᴚᴛᴜᴝ ᴞᴟᴠᴡᴢᴣᴤᴥᴬᴭᴮᴯᴰᴱᴲᴳᴴᴵᴶᴷᴸᴹᴺᴻᴼᴽᴾᴿᵀᵁᵂᵃᵄᵅᵆᵇᵈᵉᵊᵋᵌᵍᵎᵏᵐᵑᵒᵓᵔᵕᵖᵗᵘᵙᵚᵛᵜᵢᵣᵤ ᵥᵫᵬᵭᵮᵯᵰᵱᵲᵳᵴᵵᵶᵷᵹᵺᵻᵼᵽᵾᵿᶀᶁᶂᶃᶄᶅᶆᶇᶈᶉᶊᶋᶌᶍᶎᶏᶐᶑᶒᶓᶔᶕᶖᶗᶘᶙᶚᶛᶜᶝᶞᶟᶠᶡᶢᶣᶤᶥᶦ ᶧᶨᶩᶪᶫᶬᶭᶮᶯᶰᶱᶲᶳᶴᶵᶶᶷᶸᶹᶺᶻᶼᶽᶾḁḃḅḇḉḋḍḏḑḓḕḗḙḛḝḟḡḣḥḧḩḫḭḯḱḳḵḷḹḻḽḿṁṃṅṇ ṉṋṍṏṑṓṕṗṙṛṝṟṡṣṥṧṩṫṭṯṱṳṵṷṹṻṽṿẁẃẅẇẉẋẍẏẑẓẕẖẗẘẙẚẛẜẝẟạảấầẩẫậắằẳẵặẹ ẻẽếềểễệỉịọỏốồổỗộớờởỡợụủứừửữựỳỵỷỹỻỽỿⁱⁿₐₑₒₓₔₕₖₗₘₙₚₛₜⅎↄⱡⱥⱦⱨⱪ ⱬⱱⱳⱴⱶⱷⱸⱹⱺⱻⱼⱽꜣꜥꜧꜩꜫꜭꜯꜰꜱꜳꜵꜷꜹꜻꜽꜿꝁꝃꝅꝇꝉꝋꝍꝏꝑꝓꝕꝗꝙꝛꝝꝟꝡꝣꝥꝧꝩꝫꝭꝯꝰꝱꝲꝳꝴꝵ ꝶꝷꝸꝺꝼꝿꞁꞃꞅꞇꞌꞎꞑꞓꞡꞣꞥꞧꞩꟸꟹꟺfffiflffifflſtstabcdefghijklmnopqr stuvwxyz

How do you condense that again into language-specific subsets? Every letter that appears in a word in a dictionary? Then at least é belongs to German as well, even though it's usually not considered part of the German Latin subset. Unicode stays clear of that issue by simply not defining what script subsets a character belongs to (rightfully so, IMHO).

Re: How a comment on Hacker News led to 4½ new Unicode characters

#397
post #370
post #356

I got a change into Unicode 9.0 too! It was just a tweak to emoji characters to mark them all as East Asian Full Width instead of Narrow or Ambiguous so that they displayed correctly when using a fixed width font in a terminal console. This probably only matters if you like to use emoji filenames (you mad person), but it felt like a wart so I reported it & had a short back and forth with the chair of the emoji-relate…

Holy crap I appreciate this change! I thought they'd never fix it because of compatibility. Thanks for the effort you put in. It's not that I use emoji filenames, it's that I deal with real-world natural language text all the time, including at the console. (In terms of compatibility, my text-justifying function is going to stop working correctly for the period of time between when gnome-terminal updates to Unicode 9…

If they get their unicode data from the same, OS supplied source that supplies the wcwidth() function (or be using wcwidth() themselves) a libc update should fix both, I think.

Re: How a comment on Hacker News led to 4½ new Unicode characters

#398
post #397
post #370

Earlier quoted context omitted.

Holy crap I appreciate this change! I thought they'd never fix it because of compatibility. Thanks for the effort you put in. It's not that I use emoji filenames, it's that I deal with real-world natural language text all the time, including at the console. (In terms of compatibility, my text-justifying function is going to stop working correctly for the period of time between when gnome-terminal updates to Unicode 9…

If they get their unicode data from the same, OS supplied source that supplies the wcwidth() function (or be using wcwidth() themselves) a libc update should fix both, I think.

They don't. Python's "unicodedata" module updates with the minor version of Python.

This is good, actually, because the meaning of a string operation should be consistent when run on the same version of Python.

(If only this applied to the "default encoding". The default encoding should be UTF-8, not whatever you get by asking the user's likely-misconfigured locale. As it is, you can't rely on the default encoding if you want your code to work consistently.)

Re: How a comment on Hacker News led to 4½ new Unicode characters

#399
post #146

As the story mentions regarding the off symbol (a circle), there are many visually identical code points that have different semantic meanings. But in this case, they added an additional semantic meaning to an existing code point. So which is it? Does each code point represent a visual image? A semantic meaning? Both? It depends? Something else? I've tried to decipher that on my own and only learned that the answer t…

> So which is it? Does each code point represent a visual image? Look it's pretty simple, every code point represents a semantic meaning, except for: 1. those characters who also encode the width of their visual image (U+FF00..FFEF) 2. the one that means 'unknown' (U+FFFD) 3. those characters that change their visual representation depending on their position in the word (U+FB50..U+FDFF,U+FE70..U+FEFF) 4. those that…

Why would #1 #3 #4 and #5 not apply?

"every code point represents a semantic meaning" is completely consistent with the notion that some code points e.g. have differing visual representation depending on their position in the word.

Re: How a comment on Hacker News led to 4½ new Unicode characters

#400
post #218

Earlier quoted context omitted.

"many visually identical code points" The emoji code points can be represented differently on different systems given their meaning. So it makes sense to have different emojis for different 'meanings'. The 'moon' switch here does no mean 'moon' - it means 'standby' or whatever. It may look noticeably different on different systems. Think from a design perspective: you have 5 emojis to represent 'clouds, sky, earth' e…

But that doesn't explain the inconsistency in the current case. So if a system wants to render "on" differently than "straight vertical line", that's possible. However, if "off" should be rendered differently than "circle", that's not possible. (Or only possible with out-of-band information or modifier characters which would still have to be defined)

Do note that they did not include a generic "off" symbol, they included the IEEE 1621 off symbol - which must be rendered as a circle; while on the other hand the IEEE 1621 on symbol must be rendered in a manner that is often different from just "straight vertical line" in particular regarding the corners of that line.
Post reply on HN